arXiv:2409.06744q-bio.QMcs.AI2024-09ICLR被引 32

构建蛋白大模型综合评估框架,揭示其真实能力与局限。

ProteinBench: A Holistic Evaluation of Protein Foundation Models

  • 按蛋白模态关系分类任务,建立系统化评估体系。
  • 多维度评估质量、新颖性、多样性与鲁棒性,结果可量化。
  • 开源数据集与代码,支持透明研究与持续迭代。

近年来,蛋白基础模型在三维结构预测、蛋白质设计和构象动力学等任务中取得显著进展。然而,由于缺乏统一的评估框架,这些模型的能力与局限仍不清晰。为此,我们提出ProteinBench,一个全面的评估框架,包含三方面:(i) 基于蛋白模态间关系的任务分类体系;(ii) 覆盖质量、新颖性、多样性与鲁棒性的多指标评估方法;(iii) 针对不同用户目标的深度分析。对多个蛋白基础模型的评估揭示了其当前表现与限制。为促进透明性与研究合作,我们公开发布评估数据集、代码及公共排行榜,旨在建立持续演进的标准评估体系,推动领域发展。

原文摘要 · Abstract (English)

Recent years have witnessed a surge in the development of protein foundation models, significantly improving performance in protein prediction and generative tasks ranging from 3D structure prediction and protein design to conformational dynamics. However, the capabilities and limitations associated with these models remain poorly understood due to the absence of a unified evaluation framework. To fill this gap, we introduce ProteinBench, a holistic evaluation framework designed to enhance the transparency of protein foundation models. Our approach consists of three key components: (i) A taxonomic classification of tasks that broadly encompass the main challenges in the protein domain, based on the relationships between different protein modalities; (ii) A multi-metric evaluation approach that assesses performance across four key dimensions: quality, novelty, diversity, and robustness; and (iii) In-depth analyses from various user objectives, providing a holistic view of model performance. Our comprehensive evaluation of protein foundation models reveals several key findings that shed light on their current capabilities and limitations. To promote transparency and facilitate further research, we release the evaluation dataset, code, and a public leaderboard publicly for further analysis and a general modular toolkit. We intend for ProteinBench to be a living benchmark for establishing a standardized, in-depth evaluation framework for protein foundation models, driving their development and application while fostering collaboration within the field.

蛋白模型评估框架基础模型生物计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。