Skip to main navigation Skip to search Skip to main content

Fine-tuning protein language models enhances the identification and interpretation of the transcription factors

  • Mir Tanveerul Hassan
  • , Saima Gaffar
  • , Hamza Zahid
  • , Bilal Ahmad
  • , Sang Jun Lee*
  • *Corresponding author for this work
  • Jeonbuk National University
  • Sungkyunkwan University

Research output: Contribution to journalJournal articlepeer-review

Abstract

Transcription factors (TFs) are pivotal regulators of gene expression and play essential roles in diverse cellular activities. The three-dimensional organization of the genome and transcriptional regulation are predominantly orchestrated by TFs. By recruiting the transcriptional machinery to gene enhancers or promoters, TFs can either activate or repress transcription, thereby controlling gene activity and various biological pathways. Accurate identification of TFs is vital for elucidating gene regulatory mechanisms within cells. However, experimental identification remains labor-intensive and time-consuming, highlighting the necessity for efficient computational approaches. In this study, we present a two-layer predictive framework utilizing protein language models (pLMs) via full fine-tuning and parameter-efficient fine-tuning. The initial layer robustly classifies and identifies transcription factors, while the subsequent layer predicts TFs with a binding preference for methylated DNA (TFPMs). Our approach further incorporates attention weights and protein sequence motifs to enhance interpretability and predictive capability. By leveraging attention mechanisms, we highlight biologically relevant regions of the protein sequences that contribute most strongly to the predictions. Additionally, motif analysis facilitates the identification of conserved sequence patterns that are critical for TF recognition and function. Across both TF and TFPM classification tasks, the inclusion of these features allowed our methods to consistently surpass contemporary models, as demonstrated by independent test results. To promote accessibility, we provide a free web server for TF prediction at http://milab.jbnu.ac.kr/webserver/tf.

Original languageEnglish
Article number100173
JournalArtificial Intelligence in the Life Sciences
Volume9
DOIs
StatePublished - 2026.06

Keywords

  • Full fine-tuning
  • Low-rank adaption (LoRA)
  • Parameter-efficient fine-tuning (PEFT)
  • Protein language model
  • Transcription factors

Fingerprint

Dive into the research topics of 'Fine-tuning protein language models enhances the identification and interpretation of the transcription factors'. Together they form a unique fingerprint.

Cite this