构建抗体-抗原亲和力排序基准,提升模型泛化能力
AbRank: A Benchmark Dataset and Metric-Learning Framework for Antibody-Antigen Affinity Ranking
- 将亲和力预测转为成对排序任务,增强训练鲁棒性
- 整合38万+实验数据,覆盖多样抗体抗原与实验条件
- 适合抗体药物设计、疫苗开发领域的研究人员使用
准确预测抗体-抗原(Ab-Ag)结合亲和力对治疗性抗体设计和疫苗开发至关重要,但现有模型受限于噪声实验标签、异质检测条件及在庞大序列空间中的泛化能力不足。我们提出AbRank,一个大规模基准数据集与评估框架,将亲和力预测重构为成对排序问题。AbRank汇聚来自九个异源来源的超过38万组结合实验,涵盖多样抗体、抗原及实验条件,并引入标准化数据划分,系统性地增加分布偏移,从点突变等局部扰动到新型抗体与抗原间的广泛泛化。为确保可靠监督,AbRank采用m-置信排序框架,剔除亲和力差异不显著的比较对,仅保留至少具有m倍差异的样本进行训练。作为基线,我们提出WALLE-Affinity,一种基于图结构的方法,融合蛋白质语言模型嵌入与结构信息以预测成对结合偏好。基准测试揭示当前方法在真实泛化场景下的显著局限,同时表明基于排序的训练能提升模型鲁棒性与可迁移性。总之,AbRank为机器学习模型在抗体-抗原空间中的泛化提供了坚实基础,对可扩展、结构感知的抗体药物设计具有直接应用价值。
原文摘要 · Abstract (English)
Accurate prediction of antibody-antigen (Ab-Ag) binding affinity is essential for therapeutic design and vaccine development, yet the performance of current models is limited by noisy experimental labels, heterogeneous assay conditions, and poor generalization across the vast antibody and antigen sequence space. We introduce AbRank, a large-scale benchmark and evaluation framework that reframes affinity prediction as a pairwise ranking problem. AbRank aggregates over 380,000 binding assays from nine heterogeneous sources, spanning diverse antibodies, antigens, and experimental conditions, and introduces standardized data splits that systematically increase distribution shift, from local perturbations such as point mutations to broad generalization across novel antigens and antibodies. To ensure robust supervision, AbRank defines an m-confident ranking framework by filtering out comparisons with marginal affinity differences, focusing training on pairs with at least an m-fold difference in measured binding strength. As a baseline for the benchmark, we introduce WALLE-Affinity, a graph-based approach that integrates protein language model embeddings with structural information to predict pairwise binding preferences. Our benchmarks reveal significant limitations in current methods under realistic generalization settings and demonstrate that ranking-based training improves robustness and transferability. In summary, AbRank offers a robust foundation for machine learning models to generalize across the antibody-antigen space, with direct relevance for scalable, structure-aware antibody therapeutic design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。