跳至主導覽 跳至搜尋 跳過主要內容

Supervised fine-tuning enhances unsupervised learning from 45 million amino acids in TCR and peptide sequences

  • Kewei Zhou
  • , Kai Xu
  • , Shaolong Lin
  • , Silong Zhai
  • , Huanxiang Liu
  • , Xiaojun Yao
  • Macao Polytechnic University

研究成果: Article同行評審

摘要

Motivation: T cell receptor (TCR) and peptide interactions (TPI) are one of the most important parts of T cell immunity. Experimental identification of TPI is time-consuming and labor-intensive; therefore, it is necessary to develop computational prediction method that exploit existing data to predict TPI. Results: We use huge TCR and peptide sequences to pre-train two language models (∼152M parameters), respectively, and integrate them into a sequence-based only prediction framework (i.e. RoBERTcr) with supervised fine-tuning (SFT). Visualization of amino acids embedding from pre-trained language model (PLM) shows biochemical clusters based on different properties, and our PLMs outperform existing protein language models (i.e. ESM and ProtTrans) under the same condition. RoBERTcr achieved higher performance than other state-of-the-art methods based on structures or sequences without dataset bias. The visualization of attention from our framework implies valuable spatial information that residues in TCR contacting peptides are the key to their interaction. Availability: RoBERTcr is free available at https://fca_icdb.mpu.edu.mo/robertcr/ and https://doi.org/10.5281/zenodo.18043054.

原文English
文章編號btag200
期刊Bioinformatics
42
發行號5
DOIs
出版狀態Published - 5月 2026

指紋

深入研究「Supervised fine-tuning enhances unsupervised learning from 45 million amino acids in TCR and peptide sequences」主題。共同形成了獨特的指紋。

引用此