用图不变量诊断基准数据集,判断模型是否真学到了结构信息。
Invariant-Based Diagnostics for Graph Benchmarks

- 用图不变量作为结构描述符,诊断模型是否依赖图结构。
- 26个数据集上,简单不变量模型性能媲美甚至超过主流模型。
- 适合评估模型是否真正学习了图结构,推动图基础模型发展。
图基础模型的发展受限于基准测试中节点特征与图结构的混淆,难以判断模型是真正从连接关系中学习,还是根本不需要。我们提出使用图不变量——即对排列不变、任务无关的结构描述符——作为图基准的诊断框架。结果显示:(i) 不变量比标准GNN更具表达能力;(ii) 不变量能刻画基准数据集中内外的结构异质性;(iii) 不变量可预测多任务性能;(iv) 简单的不变量模型在26个数据集上表现与Transformer和消息传递基线相当,有时更优。结果表明,表达能力并非预测性能的主要驱动因素,当结构重要时,非训练的结构代理常可匹配训练过的消息传递模型。因此我们主张,不变量基线应成为评估任务是否需要结构及模型是否捕捉结构的标准,为图基础模型提供基础支撑。
原文摘要 · Abstract (English)
Progress on graph foundation models is hindered by benchmark practices that conflate the contributions of node features and graph structure, making it hard to tell whether a model actually learns from connectivity, or whether it even needs to. We propose addressing this using graph invariants, i.e., permutation-invariant, task-agnostic structural descriptors that serve as a diagnostic framework for graph benchmarks. We show that (i) invariants are more expressive than standard GNNs, (ii) invariants characterize structural heterogeneity within and across benchmark datasets, (iii) invariants predict multi-task performance, and (iv) simple invariant-based models are competitive with, and sometimes exceed, transformer and message-passing baselines across 26 datasets. Our results suggest that expressivity is not the main driver of predictive performance, and that on tasks where structure matters, a non-trainable structural proxy often matches trained message-passing models. We thus posit that invariant baselines should become a standard for evaluating whether structure is required for a task and whether a model picks up on it, serving as a stepping stone towards graph foundation models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。