摘要
Motivation: T cell receptor (TCR) and peptide interactions (TPI) are one of the most important parts of T cell immunity. Experimental identification of TPI is time-consuming and labor-intensive; therefore, it is necessary to develop computational prediction method that exploit existing data to predict TPI. Results: We use huge TCR and peptide sequences to pre-train two language models (∼152M parameters), respectively, and integrate them into a sequence-based only prediction framework (i.e. RoBERTcr) with supervised fine-tuning (SFT). Visualization of amino acids embedding from pre-trained language model (PLM) shows biochemical clusters based on different properties, and our PLMs outperform existing protein language models (i.e. ESM and ProtTrans) under the same condition. RoBERTcr achieved higher performance than other state-of-the-art methods based on structures or sequences without dataset bias. The visualization of attention from our framework implies valuable spatial information that residues in TCR contacting peptides are the key to their interaction. Availability: RoBERTcr is free available at https://fca_icdb.mpu.edu.mo/robertcr/ and https://doi.org/10.5281/zenodo.18043054.
| 原文 | English |
|---|---|
| 文章編號 | btag200 |
| 期刊 | Bioinformatics |
| 卷 | 42 |
| 發行號 | 5 |
| DOIs | |
| 出版狀態 | Published - 5月 2026 |
指紋
深入研究「Supervised fine-tuning enhances unsupervised learning from 45 million amino acids in TCR and peptide sequences」主題。共同形成了獨特的指紋。引用此
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver