Skip to main navigation Skip to search Skip to main content

CLKAN-DBPred: Leveraging pre-trained protein language models with convolutional layer Kolmogorov-Arnold Network for predicting DNA-binding proteins

  • Kamran Arshad
  • , Muhammad Arif
  • , Dong Jun Yu*
  • *Corresponding author for this work
  • Nanjing University of Science and Technology

Research output: Contribution to journalArticlepeer-review

Abstract

DNA-binding protein prediction helps to understand the fundamental process of life activities. Existing computational methods targeting large-scale DBPs are cost-effective and fast compared to conventional biochemical experiments. However, many earlier in-silico techniques rely on shallow or handcrafted descriptors that may not fully capture contextual sequence information. Recent pre-trained biological language models (BLMs) improve sequence representation by learning rich contextual characteristics from proteins. Methods: In this study, we develop CLKAN-DBPred, an integrated framework for predicting binding activity of DBPs using a fine-tuned approach that integrates DistilProtBert, Prot-T5, and ESM2 embeddings, the RFE-F7-1195 representation, and a convolutional KAN classifier. Training used hard-label cross-entropy together with a KL-divergence term computed against smoothed targets (epsilon = 0.10; kl_weight = 1.5). Additional component-wise loss ablations, KAN pathway and spline-parameter sensitivity analyses, pure machine-learning baselines, and paired statistical tests were conducted. Results: CLKAN-DBPred achieved 85.04 Acc, 0.7009 MCC, 0.8489 F1, and 0.9075 AUC in the reported fivefold CV evaluation and 81.76 Acc, 0.6352 MCC, 0.8163 F1, and 0.8905 AUC on PDB296. It obtained the highest AUC and MCC among the compared DBP predictors. The added analyses localized most of the KAN contribution to the spline pathway and showed limited sensitivity to the tested loss and spline settings. Against a direct XGBoost baseline on RFE-F7-1195, CLKAN-DBPred achieved stronger mean CV performance; paired independent-test differences in Acc, MCC, F1, and AUC were not significant. These results support CLKAN-DBPred as a competitive sequence-based framework for DBP prediction. All data and models are available at: https://doi.org/10.5281/zenodo.18023779.

Original languageEnglish
Article number105875
Number of pages19
JournalChemometrics and Intelligent Laboratory Systems
Volume278
DOIs
Publication statusPublished - 15 Nov 2026

Keywords

  • Biological language model (BLMs)
  • Deep learning
  • DNA-binding Proteins (DBPs)
  • Drug discovery
  • Feature selection

Fingerprint

Dive into the research topics of 'CLKAN-DBPred: Leveraging pre-trained protein language models with convolutional layer Kolmogorov-Arnold Network for predicting DNA-binding proteins'. Together they form a unique fingerprint.

Cite this