新基准测试发现现有模型泛化能力有限
New Benchmarking Shows Limited Generalization Power of TCR Antigenic Epitope Prediction Models

- 构建两类全新未见数据集用于公平评估
- 现有模型在新数据上表现不佳,泛化能力差
- 为下一代预测算法提供可靠评估基础
准确的计算预测T细胞受体(TCR)抗原特异性将革新T细胞生物学研究并实现可扩展的免疫工程,但现有模型在敏感性和特异性方面仍不足以支持广泛应用。主要瓶颈在于缺乏经过严格定义、未曾见过的基准数据集,无法实现对模型性能和泛化能力的无偏评估。本文描述了两类满足此标准的互补数据集,认为它们不仅为模型评估提供了稳健框架,也为下一代TCR-抗原预测算法的发展奠定了基础。
原文摘要 · Abstract (English)
Accurate computational prediction of T cell receptor (TCR) antigen specificity would transform the study of T cell biology and enable scalable immune engineering, yet existing models lack sufficient sensitivity and specificity for broad applications. A major limitation is the absence of rigorously defined, unseen benchmark datasets that allow unbiased evaluation of model performance and generalizability. Here, we describe two complementary classes of datasets that meet this criterion and argue that they provide both a robust framework for model assessment and a foundation for next-generation TCR-antigen prediction algorithm development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。