Skip to main navigation Skip to search Skip to main content

Supervised fine-tuning enhances unsupervised learning from 45 million amino acids in TCR and peptide sequences

  • Kewei Zhou
  • , Kai Xu
  • , Shaolong Lin
  • , Silong Zhai
  • , Huanxiang Liu
  • , Xiaojun Yao
  • Macao Polytechnic University

Research output: Contribution to journalArticlepeer-review

Abstract

Motivation: T cell receptor (TCR) and peptide interactions (TPI) are one of the most important parts of T cell immunity. Experimental identification of TPI is time-consuming and labor-intensive; therefore, it is necessary to develop computational prediction method that exploit existing data to predict TPI. Results: We use huge TCR and peptide sequences to pre-train two language models (∼152M parameters), respectively, and integrate them into a sequence-based only prediction framework (i.e. RoBERTcr) with supervised fine-tuning (SFT). Visualization of amino acids embedding from pre-trained language model (PLM) shows biochemical clusters based on different properties, and our PLMs outperform existing protein language models (i.e. ESM and ProtTrans) under the same condition. RoBERTcr achieved higher performance than other state-of-the-art methods based on structures or sequences without dataset bias. The visualization of attention from our framework implies valuable spatial information that residues in TCR contacting peptides are the key to their interaction. Availability: RoBERTcr is free available at https://fca_icdb.mpu.edu.mo/robertcr/ and https://doi.org/10.5281/zenodo.18043054.

Original languageEnglish
Article numberbtag200
JournalBioinformatics
Volume42
Issue number5
DOIs
Publication statusPublished - May 2026

Fingerprint

Dive into the research topics of 'Supervised fine-tuning enhances unsupervised learning from 45 million amino acids in TCR and peptide sequences'. Together they form a unique fingerprint.

Cite this