提出评估图学习数据集质量的新框架,解决当前评测标准缺失问题。
No Metric to Rule Them All: Toward Principled Evaluations of Graph-Learning Datasets
- 通过扰动图结构和节点特征双模态,量化数据集差异。
- 引入性能可分性与模态互补性两项指标,分别检验模型区分能力与模态协同效果。
- 适用于数据构建者、评测设计者,推动更可信的图学习研究基准。
基准数据集对图学习的发展至关重要,但现有数据集存在诸多问题,例如忽略图结构的方法反而表现更优。这引发两个核心问题:何为优质图学习数据集?如何评估其质量?传统评估范式依赖数据集评测模型,无法反向评价数据集本身。为此,本文从基础出发,基于图数据特有的结构与特征双重模式,提出可扩展的环状扰动框架(Rings),通过对比原始数据与扰动版本的差异来评估数据集质量。在此框架下,定义了两项评价指标:性能可分性(衡量模型区分能力)和模态互补性(衡量双模态协同效果)。在图级别任务上广泛实验验证了该框架的有效性,并为提升图学习方法评估提供了具体建议。本工作开辟了以数据为中心的图学习新方向,迈向系统化评估评测体系的重要一步。
原文摘要 · Abstract (English)
Benchmark datasets have proved pivotal to the success of graph learning, and good benchmark datasets are crucial to guide the development of the field. Recent research has highlighted problems with graph-learning datasets and benchmarking practices -- revealing, for example, that methods which ignore the graph structure can outperform graph-based approaches. Such findings raise two questions: (1) What makes a good graph-learning dataset, and (2) how can we evaluate dataset quality in graph learning? Our work addresses these questions. As the classic evaluation setup uses datasets to evaluate models, it does not apply to dataset evaluation. Hence, we start from first principles. Observing that graph-learning datasets uniquely combine two modes -- graph structure and node features --, we introduce Rings, a flexible and extensible mode-perturbation framework to assess the quality of graph-learning datasets based on dataset ablations -- i.e., quantifying differences between the original dataset and its perturbed representations. Within this framework, we propose two measures -- performance separability and mode complementarity -- as evaluation tools, each assessing the capacity of a graph dataset to benchmark the power and efficacy of graph-learning methods from a distinct angle. We demonstrate the utility of our framework for dataset evaluation via extensive experiments on graph-level tasks and derive actionable recommendations for improving the evaluation of graph-learning methods. Our work opens new research directions in data-centric graph learning, and it constitutes a step toward the systematic evaluation of evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。