ImmSET模型可高效预测T细胞受体与抗原的结合特异性,助力个性化免疫治疗。
ImmSET: Sequence-Based Predictor of TCR-pMHC Specificity at Scale
- 基于序列的Transformer架构,直接建模多变长生物序列间的相互作用。
- 在大规模数据下性能随数据量持续提升,优于微调的ESM2和AlphaFold管道。
- 解决旧方法的评估偏差问题,适合高通量免疫数据研究与药物设计者使用。
T细胞是适应性免疫系统的关键组成部分,在传染病、自身免疫病和癌症中起重要作用。T细胞功能由T细胞受体(TCR)介导,这是一种高度多样化的受体,可靶向由主要组织相容性复合体(pMHC)呈递的特定肽段。准确预测TCR与其对应pMHC的特异性,对理解适应性免疫机制及实现个性化治疗至关重要。然而,由于TCR和pMHC均存在极端多样性,该蛋白-蛋白相互作用的预测仍具挑战性。本文提出ImmSET(Immune Synapse Encoding Transformer),一种新型序列基础架构,用于建模可变长度生物序列集合间的相互作用。我们在多种数据集规模和构成下训练该模型,并研究其在不同pMHC靶标上的泛化能力。我们揭示了以往序列方法中存在的失败模式,该模式会夸大其在任务中的表现,而ImmSET在更严格评估下仍保持稳健。通过系统测试ImmSET在训练数据量下的扩展行为,我们发现其性能随数据量持续提升,且在多个数据类型上优于在相同数据集上微调的预训练蛋白语言模型ESM2。最后,当提供足够训练数据时,ImmSET的表现超越AlphaFold2和AlphaFold3相关流程。本工作确立了ImmSET作为多序列互作问题的可扩展建模范式,已在TCR-pMHC场景中验证,但可推广至其他需要高通量序列推理的生物领域。
原文摘要 · Abstract (English)
T cells are a critical component of the adaptive immune system, playing a role in infectious disease, autoimmunity, and cancer. T cell function is mediated by the T cell receptor (TCR) protein, a highly diverse receptor targeting specific peptides presented by the major histocompatibility complex (pMHCs). Predicting the specificity of TCRs for their cognate pMHCs is central to understanding adaptive immunity and enabling personalized therapies. However, accurate prediction of this protein-protein interaction remains challenging due to the extreme diversity of both TCRs and pMHCs. Here, we present ImmSET (Immune Synapse Encoding Transformer), a novel sequence-based architecture designed to model interactions among sets of variable-length biological sequences. We train this model across a range of dataset sizes and compositions and study the resulting models' generalization to pMHC targets. We describe a failure mode in prior sequence-based approaches that inflates previously reported performance on this task and show that ImmSET remains robust under stricter evaluation. In systematically testing the scaling behavior of ImmSET with training data, we show that performance scales consistently with data volume across multiple data types and compares favorably with the pre-trained protein language model ESM2 fine-tuned on the same datasets. Finally, we demonstrate that ImmSET can outperform AlphaFold2 and AlphaFold3-based pipelines on TCR-pMHC specificity prediction when provided sufficient training data. This work establishes ImmSET as a scalable modeling paradigm for multi-sequence interaction problems, demonstrated in the TCR-pMHC setting but generalizable to other biological domains where high-throughput sequence-driven reasoning complements structure prediction and experimental mapping.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。