提出可解释的结构正则化模型,提升T细胞受体与抗原结合预测性能。
Structure-Regularized Interpretable TCR-Epitope Prediction

- 融合语言模型与接触原型,实现残基级相互作用建模
- 在TCR-XAI基准上达到当前最佳性能,解释质量更优
- 揭示生成结构对模型学习的影响,适合免疫机制研究者
T细胞受体(TCR)-抗原表位结合预测对于理解适应性免疫和开发免疫疗法至关重要。现有基于序列或结构的模型往往在未见表位上泛化能力差,且可解释性有限。此外,生成结构对模型学习的影响尚不明确。我们提出TCR-SRIM,一种结构正则化的可解释设计模型,结合蛋白质语言模型嵌入与可解释接触原型,捕捉残基级的TCR-表位相互作用。TCR-SRIM在TCR-XAI基准上实现最先进的预测性能和更好的解释质量。利用其内在可解释性,进一步评估生成结构对模型学习的影响。尽管AlphaFold3、TCRModel2和tFold-TCR生成的结构表现良好,但其相互作用模式准确性较低,结合位点多样性不足,相比实验解析结构仍有差距。结果表明当前结构预测模型在TCR-表位学习中的局限性,并凸显可解释设计模型在研究生成生物结构中的价值。
原文摘要 · Abstract (English)
T cell receptor (TCR)-epitope binding prediction is essential for understanding adaptive immunity and developing immunotherapies. Existing sequence- and structure-based models often generalize poorly to unseen epitopes and provide limited interpretability. Furthermore, the impact of generated structures on model learning remains unclear. We present TCR-SRIM, a structure-regularized interpretable-by-design model that combines protein language model embeddings with interpretable contact prototypes to capture residue-level TCR-epitope interactions. TCR-SRIM achieves state-of-the-art predictive performance and improved interpretation quality on the TCR-XAI benchmark. Using its inherent interpretability, we further evaluate the effect of generated structures on model learning. While structures predicted by AlphaFold3, TCRModel2, and tFold-TCR yield competitive performance, they lead to less accurate interaction patterns and reduced binding-site diversity than experimentally-resolved structures. Our results highlight limitations of current structure prediction models for TCR-epitope learning and demonstrate the value of interpretable-by-design models for studying generated biological structures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。