首个针对功能导向蛋白设计的综合性评估基准,解决模型评测不统一问题。
PDFBench: A Benchmark for De novo Protein Design from Function
- 构建涵盖16项指标的评测框架,覆盖描述与关键词两种设计场景。
- 引入新数据集SwissTest并重用Mol-Instructions数据,确保评测数据可靠。
- 揭示不同评价指标间关联,为未来研究提供方向指导。
功能导向的蛋白质设计在药物发现和酶工程中具有重要意义,但当前缺乏统一、全面的评估框架。现有模型使用不一致且有限的评估指标,难以进行公平比较,也难以理解各评价标准间的关联。为此,我们提出PDFBench,首个面向功能引导的从头蛋白设计的综合性基准。该基准系统性地评估了八种前沿模型,在两个关键场景下使用16项指标:描述引导设计(复用原无量化评估的Mol-Instructions数据集)与关键词引导设计(引入新测试集SwissTest,采用严格时间截断以保障数据完整性)。通过跨多指标评测及相关性分析,PDFBench实现了更可靠的模型对比,并为未来研究提供了关键洞见。
原文摘要 · Abstract (English)
Function-guided protein design is a crucial task with significant applications in drug discovery and enzyme engineering. However, the field lacks a unified and comprehensive evaluation framework. Current models are assessed using inconsistent and limited subsets of metrics, which prevents fair comparison and a clear understanding of the relationships between different evaluation criteria. To address this gap, we introduce PDFBench, the first comprehensive benchmark for function-guided denovo protein design. Our benchmark systematically evaluates eight state-of-the-art models on 16 metrics across two key settings: description-guided design, for which we repurpose the Mol-Instructions dataset, originally lacking quantitative benchmarking, and keyword-guided design, for which we introduce a new test set, SwissTest, created with a strict datetime cutoff to ensure data integrity. By benchmarking across a wide array of metrics and analyzing their correlations, PDFBench enables more reliable model comparisons and provides key insights to guide future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。