arXiv:2512.04475cs.LGcs.AI2025-12被引 9

构建跨领域图学习新基准,统一评估标准

GraphBench: Next-generation graph learning benchmarking

  • 覆盖节点、边、图及生成任务的综合评测体系
  • 提供一致数据划分与分布外泛化评估协议
  • 适配图神经网络与图变压器模型的基线测试

图机器学习在分子性质预测、芯片设计等领域取得显著进展,但基准评测仍碎片化,依赖狭窄的任务特定数据集和不一致的评估协议,影响可复现性与整体进步。随着图基础模型兴起,现有基准已显不足。为此,我们提出GraphBench,一个涵盖多样真实世界领域和任务场景的综合性基准套件,包含节点级、边级、图级及生成任务。该套件提供标准化评估协议,包括一致的数据集划分、评估指标,用于衡量选定任务的分布外泛化能力,并配备统一的超参数调优框架。我们进一步用最新的消息传递神经网络与图变压器模型对GraphBench进行评估,为未来研究建立合理基线。详情见www.graphbench.io。

原文摘要 · Abstract (English)

Machine learning on graphs has made substantial progress across domains such as molecular property prediction and chip design. Yet benchmarking practices remain fragmented, often relying on narrow, task-specific datasets and inconsistent evaluation protocols, hindering reproducibility and broader progress. With the recent popularity of graph foundation models, these weaknesses have become apparent, as existing benchmarks are insufficient for thorough evaluation. To address these challenges, we introduce GraphBench, a comprehensive benchmark suite spanning diverse real-world domains and task settings, including node-level, edge-level, graph-level, and generative tasks. GraphBench provides standardized evaluation protocols, including consistent dataset splits and metrics for assessing out-of-distribution generalization across selected tasks, as well as a unified hyperparameter-tuning framework. We further evaluate GraphBench with recent message-passing neural networks and graph transformer models, establishing principled baselines for future research. See www.graphbench.io for further details.

图学习基准评测GNN

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。