arXiv:2409.05755cs.LG2024-09被引 5

重新评估异质图学习中模型与度量的真正挑战。

Revealing the Pitfalls and Re-Evaluating the Advancement of Heterophilic Graph Learning

  • 构建三类异质图数据集分类,区分真正难的和误导性的任务。
  • 在27个基准数据集上微调11种SOTA模型,发现多数方法性能不佳。
  • 首次对11种同质性度量进行量化评估,揭示其不可靠性。

过去十年,图神经网络(GNNs)在关系型数据上取得了显著进展。然而,近期研究发现异质性会导致GNN在节点级任务上性能严重下降。尽管已有大量异质性基准数据集和同质性度量被提出,但评估仍存在多重陷阱:缺乏超参数调优、对真正挑战性数据集评估不足、合成图上同质性度量缺乏定量验证。为此,我们在27个常用基准数据集上训练并微调基线模型,将其分为恶性、良性与模糊三类异质图数据集。我们首次提出该分类体系,识别出恶性与模糊异质性为真正困难的任务子集。随后,在不同类别上用微调后的超参数重新评估11种SOTA GNNs,涵盖六种主流方法,全面重评其在异质性上的有效性。最后,我们对11种流行同质性度量在三种不同生成方式的合成图上进行定量评估,首次提供详细分析,揭示基于观察的比较存在不可靠性。

原文摘要 · Abstract (English)

Over the past decade, Graph Neural Networks (GNNs) have achieved great success on machine learning tasks with relational data. However, recent studies have found that heterophily can cause significant performance degradation of GNNs, especially on node-level tasks. Numerous heterophilic benchmark datasets have been put forward to validate the efficacy of heterophily-specific GNNs, and various homophily metrics have been designed to help recognize these challenging datasets. Nevertheless, there still exist multiple pitfalls that severely hinder the proper evaluation of new models and metrics: 1) lack of hyperparameter tuning; 2) insufficient evaluation on the truly challenging heterophilic datasets; 3) missing quantitative evaluation for homophily metrics on synthetic graphs. To overcome these challenges, we first train and fine-tune baseline models on $27$ most widely used benchmark datasets, and categorize them into three distinct groups: malignant, benign and ambiguous heterophilic datasets. We identify malignant and ambiguous heterophily as the truly challenging subsets of tasks, and to our best knowledge, we are the first to propose such taxonomy. Then, we re-evaluate $11$ state-of-the-arts (SOTA) GNNs, covering six popular methods, with fine-tuned hyperparameters on different groups of heterophilic datasets. Based on the model performance, we comprehensively reassess the effectiveness of different methods on heterophily. At last, we evaluate $11$ popular homophily metrics on synthetic graphs with three different graph generation approaches. To overcome the unreliability of observation-based comparison and evaluation, we conduct the first quantitative evaluation and provide detailed analysis.

异质图GNN评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。