arXiv:2507.05101cs.LGcs.AI2025-07NeurIPS被引 4

首个从图角度评估蛋白质互作的基准,推动真实生物应用

PRING: Rethinking Protein-Protein Interaction Prediction from Pairs to Graphs

  • 将蛋白互作预测从成对评估升级为网络图级,关注整体结构与功能
  • 构建含21,484蛋白、186,818互作的多物种高质量数据集,防冗余和泄露
  • 适合关注生物网络建模、功能注释与疾病机制研究的科研人员

基于深度学习的计算方法在预测蛋白质-蛋白质互作(PPI)方面已取得显著进展。然而,现有基准主要聚焦于孤立的成对评估,忽视了模型重建具有生物学意义的PPI网络的能力,而这对于生物研究至关重要。为此,我们提出PRING,首个从图级别评估PPI预测的综合性基准。PRING构建了一个高质量、多物种的PPI网络数据集,包含21,484个蛋白和186,818个互作,并采用精心设计的策略解决数据冗余与泄露问题。在此金标准数据集基础上,我们建立了两种互补的评估范式:(1) 以拓扑为导向的任务,评估跨物种及物种内PPI网络的构建能力;(2) 以功能为导向的任务,包括蛋白质复合物通路预测、GO模块分析与关键蛋白验证。这些评估不仅反映模型对网络拓扑的理解能力,也支持蛋白功能注释、生物模块识别乃至疾病机制分析。在四类代表性模型(基于序列相似性、朴素序列、蛋白语言模型及结构的方法)上的大量实验表明,当前PPI模型在恢复PPI网络的结构与功能特性方面仍存在明显局限,凸显其在支持真实生物应用方面的差距。我们相信PRING为社区提供了可靠的平台,以推动更高效PPI预测模型的发展。PRING的数据集与源代码可在https://github.com/SophieSarceau/PRING获取。

原文摘要 · Abstract (English)

Deep learning-based computational methods have achieved promising results in predicting protein-protein interactions (PPIs). However, existing benchmarks predominantly focus on isolated pairwise evaluations, overlooking a model's capability to reconstruct biologically meaningful PPI networks, which is crucial for biology research. To address this gap, we introduce PRING, the first comprehensive benchmark that evaluates protein-protein interaction prediction from a graph-level perspective. PRING curates a high-quality, multi-species PPI network dataset comprising 21,484 proteins and 186,818 interactions, with well-designed strategies to address both data redundancy and leakage. Building on this golden-standard dataset, we establish two complementary evaluation paradigms: (1) topology-oriented tasks, which assess intra and cross-species PPI network construction, and (2) function-oriented tasks, including protein complex pathway prediction, GO module analysis, and essential protein justification. These evaluations not only reflect the model's capability to understand the network topology but also facilitate protein function annotation, biological module detection, and even disease mechanism analysis. Extensive experiments on four representative model categories, consisting of sequence similarity-based, naive sequence-based, protein language model-based, and structure-based approaches, demonstrate that current PPI models have potential limitations in recovering both structural and functional properties of PPI networks, highlighting the gap in supporting real-world biological applications. We believe PRING provides a reliable platform to guide the development of more effective PPI prediction models for the community. The dataset and source code of PRING are available at https://github.com/SophieSarceau/PRING.

蛋白质互作图神经网络生物网络基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。