建立标准化蛋白结合物设计评估框架,解决实验可比性难题。
ProtDBench: A Unified Benchmark of Protein Binder Design and Evaluation

- 统一评测任务与协议,避免不同研究间指标不可比
- 在10个靶点上测试开源生成方法,发现验证器依赖偏差显著
- 引入时效与结构多样性指标,更贴近真实研发场景
近期从头设计蛋白结合物的研究不断取得实验验证进展,但现有计算评估指标因缺乏统一标准,难以跨研究比较。本文提出ProtDBench,一个标准化且考虑吞吐量的蛋白结合物设计评估框架。该框架定义了统一的任务、评估协议和成功标准,可系统分析评估设计对性能结果的影响。基于大规模湿实验标注数据集,我们分析了常用结构预测模型作为验证器时的表现,发现其存在显著的验证器依赖偏差,且在相同过滤条件下一致性有限。随后,在固定评估协议下,对十种不同蛋白靶点上的代表性开源生成方法进行了基准测试。除序列成功率外,还引入基于24小时计算预算的吞吐量感知指标,以及考虑结构多样性的聚类级成功标准。结果揭示了过滤规则、成功定义与吞吐量评估之间的系统性差异,影响计算效率、成功率与结构多样性。总体而言,ProtDBench提供了一个公平、可复现的评估流程,支持在真实评估设定下对蛋白结合物设计方法进行系统化、可控的比较。
原文摘要 · Abstract (English)
Recent advances in de novo protein binder design have enabled increasing experimental validation, yet reported in silico metrics remain difficult to interpret or compare across studies due to non-standardized evaluation protocols. We introduce ProtDBench, a standardized and throughput-aware evaluation framework for protein binder design. ProtDBench defines unified benchmark tasks, evaluation protocols, and success criteria, enabling systematic analysis of how evaluation design influences observed performance. Using a large wet-lab annotated dataset, we analyze commonly used structure prediction models as evaluation verifiers, revealing substantial verifier-dependent bias and limited agreement under identical filtering protocols. We then benchmark representative open-source generative binder design methods across ten diverse protein targets under a fixed evaluation protocol. Beyond per-sequence success rates, ProtDBench incorporates throughput-aware metrics based on a fixed 24-hour budget, as well as cluster-level success criteria to account for structural diversity. Together, these results expose systematic differences induced by filtering rules, success definitions, and throughput-aware evaluation between computational efficiency, success rate, and structural diversity. Overall, ProtDBench provides a fair and reproducible evaluation pipeline that supports systematic and controlled comparison of protein binder design methods under realistic evaluation settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。