arXiv:2505.22674q-bio.BMcs.AI2025-05NeurIPS被引 1

构建百万级蛋白复合物结构模型评估基准,助力精准预测与药物研发。

PSBench: a large-scale benchmark for estimating the accuracy of protein complex structural models

  • 基于CASP15/16竞赛数据构建多层级标注的百万模型基准集
  • 提出GATE图神经网络方法,在盲测中跻身顶级评估模型行列
  • 适合蛋白质结构预测、药物设计及机器学习研究者使用

预测蛋白复合物结构对功能分析、蛋白设计和药物发现至关重要。尽管AlphaFold等AI方法能准确预测多数蛋白复合物结构,但可靠评估预测模型质量(模型精度估计,EMA)仍面临挑战。主要瓶颈在于缺乏大规模、多样化且标注完善的训练与评估数据。为此,我们推出PSBench,一个包含四套大规模标注数据集的基准套件,数据源自第15和第16届社区级蛋白质结构预测评估竞赛(CASP15和CASP16)。PSBench涵盖超过一百万种结构模型,覆盖广泛的蛋白序列长度、复合物组成、功能类别和建模难度。每个模型均在全局、局部和界面层面标注多种互补的质量评分。该基准还提供多种评估指标和基线EMA方法,支持严格比较。为验证其价值,我们在CASP15数据上训练并评估了基于图变换器的GATE EMA方法,该方法在2024年CASP16盲测中表现优异,位列顶尖。这些结果表明PSBench是推动蛋白复合物建模中EMA研究的重要资源。项目已开源:https://github.com/BioinfoMachineLearning/PSBench。

原文摘要 · Abstract (English)

Predicting protein complex structures is essential for protein function analysis, protein design, and drug discovery. While AI methods like AlphaFold can predict accurate structural models for many protein complexes, reliably estimating the quality of these predicted models (estimation of model accuracy, or EMA) for model ranking and selection remains a major challenge. A key barrier to developing effective machine learning-based EMA methods is the lack of large, diverse, and well-annotated datasets for training and evaluation. To address this gap, we introduce PSBench, a benchmark suite comprising four large-scale, labeled datasets generated during the 15th and 16th community-wide Critical Assessment of Protein Structure Prediction (CASP15 and CASP16). PSBench includes over one million structural models covering a wide range of protein sequence lengths, complex stoichiometries, functional classes, and modeling difficulties. Each model is annotated with multiple complementary quality scores at the global, local, and interface levels. PSBench also provides multiple evaluation metrics and baseline EMA methods to facilitate rigorous comparisons. To demonstrate PSBench's utility, we trained and evaluated GATE, a graph transformer-based EMA method, on the CASP15 data. GATE was blindly tested in CASP16 (2024), where it ranked among the top-performing EMA methods. These results highlight PSBench as a valuable resource for advancing EMA research in protein complex modeling. PSBench is publicly available at: https://github.com/BioinfoMachineLearning/PSBench.

蛋白结构机器学习评估基准生物信息

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。