统一处理同质与异质文本图,提升图基础模型泛化能力
H$^2$GFM: Towards unifying Homogeneity and Heterogeneity on Text-Attributed Graphs
- 通过统一文本空间建模跨图元关系,融合上下文语义
- 在多种同质/异质图上实现更优节点表征,提升下游任务性能
- 适合需要跨图类型泛化的图学习研究者和应用开发者
图学习在多个领域的广泛应用推动了图基础模型(GFM)的发展,旨在实现跨图与任务的通用性。现有研究主要基于具有文本属性的同质图(HoTAGs)来应对图间节点特征的异质性,但对包含多种节点/边类型的异质图(HeTAGs)探索不足。为此,本文提出H²GFM框架,可统一处理HoTAGs与HeTAGs。模型将不同图中的元关系投影至统一文本空间,并利用上下文编码捕捉空间与高阶语义关系。为获得鲁棒节点表示,提出上下文自适应图变压器(CGT),有效整合上下文邻居及其关系信息;同时引入多专家CGT结构以捕获不同图类型间的结构模式异质性。在广泛范围的HoTAGs与HeTAGs及学习场景上的实验验证了模型的有效性。
原文摘要 · Abstract (English)
The growing interests and applications of graph learning in diverse domains have propelled the development of a unified model generalizing well across different graphs and tasks, known as the Graph Foundation Model (GFM). Existing research has leveraged text-attributed graphs (TAGs) to tackle the heterogeneity in node features among graphs. However, they primarily focus on homogeneous TAGs (HoTAGs), leaving heterogeneous TAGs (HeTAGs), where multiple types of nodes/edges reside, underexplored. To enhance the capabilities and applications of GFM, we introduce H$^2$GFM, a novel framework designed to generalize across both HoTAGs and HeTAGs. Our model projects diverse meta-relations among graphs under a unified textual space, and employs a context encoding to capture spatial and higher-order semantic relationships. To achieve robust node representations, we propose a novel context-adaptive graph transformer (CGT), effectively capturing information from both context neighbors and their relationships. Furthermore, we employ a mixture of CGT experts to capture the heterogeneity in structural patterns among graph types. Comprehensive experiments on a wide range of HoTAGs and HeTAGs as well as learning scenarios demonstrate the effectiveness of our model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。