arXiv:2506.14796q-bio.BMcs.AI2025-06被引 8

构建38项任务的蛋白基础模型评测基准,揭示模型性能与任务关联。

PFMBench: Protein Foundation Model Benchmark

  • 设计跨8个领域的38项任务,统一评估17个先进模型。
  • 发现任务间存在内在关联,识别出表现最优的模型。
  • 为新模型开发提供可复现的评测流程,适合蛋白研究者使用。

本研究探讨了蛋白基础模型研究的现状与未来方向。尽管近年来进展显著,但该领域缺乏全面的评测基准,难以公平评估与深入理解模型性能。自ESM-1B以来,众多蛋白基础模型涌现,各具独特数据集与方法,但评估常局限于特定任务,限制了对泛化能力与局限性的洞察。研究人员难以厘清任务间关系、评估模型跨任务表现,也缺乏新模型开发的标准。为此,我们提出PFMBench,一个涵盖38项任务、覆盖8个关键领域的综合性评测基准。通过对17个前沿模型在38个任务上进行数百次实验,PFMBench揭示了任务间的内在关联,识别出领先模型,并提供标准化评估协议。代码已开源于GitHub。

原文摘要 · Abstract (English)

This study investigates the current landscape and future directions of protein foundation model research. While recent advancements have transformed protein science and engineering, the field lacks a comprehensive benchmark for fair evaluation and in-depth understanding. Since ESM-1B, numerous protein foundation models have emerged, each with unique datasets and methodologies. However, evaluations often focus on limited tasks tailored to specific models, hindering insights into broader generalization and limitations. Specifically, researchers struggle to understand the relationships between tasks, assess how well current models perform across them, and determine the criteria in developing new foundation models. To fill this gap, we present PFMBench, a comprehensive benchmark evaluating protein foundation models across 38 tasks spanning 8 key areas of protein science. Through hundreds of experiments on 17 state-of-the-art models across 38 tasks, PFMBench reveals the inherent correlations between tasks, identifies top-performing models, and provides a streamlined evaluation protocol. Code is available at \href{https://github.com/biomap-research/PFMBench}{\textcolor{blue}{GitHub}}.

蛋白模型评测基准基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。