用语言模型提升免疫预测,尤其在数据少时表现更优
Enhancing TCR-Peptide Interaction Prediction with Pretrained Language Models and Molecular Representations
- 结合蛋白语言模型与分子表示,捕捉免疫识别关键特征
- 零样本和少量样本下性能超越现有模型,准确率显著提升
- 适合免疫治疗、疫苗研发人员,助力个性化疗法开发
理解T细胞受体(TCRs)与肽-主要组织相容性复合物(pMHCs)之间的结合特异性,对免疫治疗和疫苗开发至关重要。然而,现有预测模型在数据稀缺和面对新表位时泛化能力不足。我们提出LANTERN(大型语言模型驱动的TCR增强识别网络),一个深度融合大规模蛋白语言模型与肽的化学表示的深度学习框架。通过使用ESM-1b编码TCR β链序列,并将肽序列转换为SMILES字符串,由MolFormer处理,LANTERN捕获了对TCR-肽识别至关重要的丰富生物与化学特征。在与ChemBERTa、TITAN和NetTCR等模型的广泛对比中,LANTERN在零样本和少样本学习场景下表现出更优性能。模型还采用稳健的负采样策略,嵌入分析显示聚类效果显著改善。这些结果表明,LANTERN有望推动TCR-pMHC结合预测的发展,支持个性化免疫治疗的实现。
原文摘要 · Abstract (English)
Understanding the binding specificity between T-cell receptors (TCRs) and peptide-major histocompatibility complexes (pMHCs) is central to immunotherapy and vaccine development. However, current predictive models struggle with generalization, especially in data-scarce settings and when faced with novel epitopes. We present LANTERN (Large lAnguage model-powered TCR-Enhanced Recognition Network), a deep learning framework that combines large-scale protein language models with chemical representations of peptides. By encoding TCR \b{eta}-chain sequences using ESM-1b and transforming peptide sequences into SMILES strings processed by MolFormer, LANTERN captures rich biological and chemical features critical for TCR-peptide recognition. Through extensive benchmarking against existing models such as ChemBERTa, TITAN, and NetTCR, LANTERN demonstrates superior performance, particularly in zero-shot and few-shot learning scenarios. Our model also benefits from a robust negative sampling strategy and shows significant clustering improvements via embedding analysis. These results highlight the potential of LANTERN to advance TCR-pMHC binding prediction and support the development of personalized immunotherapies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。