arXiv:2410.14961cs.LGcs.AI2024-10被引 14

用大模型做图模型,不靠专用模块也能通吃各类图任务

LangGFM: A Large Language Model Alone Can be a Powerful Graph Foundation Model

  • 全靠大语言模型实现图基础模型,无专用模块
  • 在26个数据集上性能达或超越当前最好水平
  • 为图学习提供统一评测基准,适合研究通用模型者

图基础模型(GFMs)近年备受关注,但不同研究采用的处理与评估方式不一致,阻碍了对进展的深入理解。现有工作多聚焦特定图学习任务(如结构、节点级或分类任务),常引入针对性模块,削弱了对其他任务的适用性,违背了基础模型应具备通用性的初衷。为此,我们提出GFMBench——一个系统且全面的基准,包含26个数据集。同时引入LangGFM,一种完全依赖大语言模型的新型图基础模型。通过重新审视有效的图文本化原则,并将图增强与自监督学习中的成功技术迁移至语言空间,LangGFM在GFMBench上表现达到或超过当前最佳水平,为图基础模型的发展提供了新视角、新经验与新基线。

原文摘要 · Abstract (English)

Graph foundation models (GFMs) have recently gained significant attention. However, the unique data processing and evaluation setups employed by different studies hinder a deeper understanding of their progress. Additionally, current research tends to focus on specific subsets of graph learning tasks, such as structural tasks, node-level tasks, or classification tasks. As a result, they often incorporate specialized modules tailored to particular task types, losing their applicability to other graph learning tasks and contradicting the original intent of foundation models to be universal. Therefore, to enhance consistency, coverage, and diversity across domains, tasks, and research interests within the graph learning community in the evaluation of GFMs, we propose GFMBench-a systematic and comprehensive benchmark comprising 26 datasets. Moreover, we introduce LangGFM, a novel GFM that relies entirely on large language models. By revisiting and exploring the effective graph textualization principles, as well as repurposing successful techniques from graph augmentation and graph self-supervised learning within the language space, LangGFM achieves performance on par with or exceeding the state of the art across GFMBench, which can offer us new perspectives, experiences, and baselines to drive forward the evolution of GFMs.

图神经网络大模型自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。