arXiv:2602.04768cs.LGcs.AI2026-02被引 3

构建百亿参数图神经网络,实现大规模异构图的通用预训练与高效适配。

Billion-Scale Graph Foundation Models

  • 提出GraphBFF Transformer架构,支持百亿参数图模型的端到端训练。
  • 在十项真实任务上性能超越基线,少样本设置下最高提升31 PRAUC点。
  • 提供数据批处理、预训练与微调方法,推动工业级图学习基础模型落地。

图结构数据支撑众多关键应用。尽管基础模型通过大规模预训练和轻量适配已改变语言与视觉领域,但将其范式扩展至通用真实世界图仍具挑战。本文提出Graph Billion-Foundation-Fusion(GraphBFF):一种构建百亿参数图基础模型(GFMs)的端到端方案,适用于大规模异构图。核心是GraphBFF Transformer,一种灵活可扩展的架构,专为实际百亿规模的图基础模型设计。基于该框架,我们揭示了异构图上的神经缩放规律,表明损失随模型容量或训练数据增长而可预测地下降,取决于当前瓶颈。GraphBFF提供具体的数据批处理、预训练与微调方法论。我们在一个真实世界的百亿规模图上验证了框架有效性,采用所提方案训练的百亿参数GraphBFF Transformer在十项不同下游任务上表现优异,涵盖节点与边级分类与回归任务,且在未见图上一致优于基线,最大提升达31 PRAUC点,包括少样本场景。最后,讨论了实现工业级图基础模型的关键挑战与开放机遇。

原文摘要 · Abstract (English)

Graph-structured data underpins many critical applications. While foundation models have transformed language and vision via large-scale pretraining and lightweight adaptation, extending this paradigm to general, real-world graphs is challenging. In this work, we present Graph Billion-Foundation-Fusion (GraphBFF): an end-to-end recipe for building billion-parameter Graph Foundation Models (GFMs) for large-scale heterogeneous graphs. Central to the recipe is the GraphBFF Transformer, a flexible and scalable architecture designed for practical billion-scale GFMs. Using the GraphBFF, we present neural scaling laws for heterogeneous graphs and show that loss decreases predictably as either model capacity or training data scales, depending on which factor is the bottleneck. The GraphBFF framework provides concrete methodologies for data batching, pretraining, and fine-tuning for building GFMs at scale. We demonstrate the effectiveness of the framework over a real-world billion-scale graph, with an evaluation of a billion-parameter GraphBFF Transformer following the proposed recipe. Across ten diverse, real-world downstream tasks on graphs unseen during training, spanning node- and link-level classification and regression, GraphBFF consistently outperforms baselines, with large margins of up to 31 PRAUC points, including in few-shot settings. Finally, we discuss key challenges and open opportunities for making GFMs a practical and principled foundation for graph learning at industrial scale.

图神经网络大模型预训练异构图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。