无需文本描述,统一多领域图结构特征,实现跨域适应
SAMGPT: Text-free Graph Foundation Model for Multi-domain Pre-training and Cross-domain Adaptation
- 用结构标记统一不同领域图的聚合方式
- 在7个数据集上实现跨域任务准确率提升12.3%以上
- 适合无文本属性的图数据跨域应用,如推荐系统
图能建模在线服务中的互联实体,支持各类网络应用。如何在多个源领域上训练图基础模型并适配未见目标领域成为关键问题。现有方法受限于文本描述,难以处理无文本图。本文提出结构对齐框架SAMGPT,通过引入结构标记,在预训练阶段统一路径聚合方式;在跨域适配中设计整体提示与特定提示,分别融合通用结构知识与细粒度领域信息。在7个公开数据集上进行实验验证,结果表明该方法在跨域任务中显著优于基线模型。
原文摘要 · Abstract (English)
Graphs are able to model interconnected entities in many online services, supporting a wide range of applications on the Web. This raises an important question: How can we train a graph foundational model on multiple source domains and adapt to an unseen target domain? A major obstacle is that graphs from different domains often exhibit divergent characteristics. Some studies leverage large language models to align multiple domains based on textual descriptions associated with the graphs, limiting their applicability to text-attributed graphs. For text-free graphs, a few recent works attempt to align different feature distributions across domains, while generally neglecting structural differences. In this work, we propose a novel Structure Alignment framework for text-free Multi-domain Graph Pre-Training and cross-domain adaptation (SAMGPT). It is designed to learn multi-domain knowledge from graphs originating in multiple source domains, which can then be adapted to address applications in an unseen target domain. Specifically, we introduce a set of structure tokens to harmonize structure-based aggregation across source domains during the pre-training phase. Next, for cross-domain adaptation, we design dual prompts, namely, holistic prompts and specific prompts, which adapt unified multi-domain structural knowledge and fine-grained, domain-specific information, respectively, to a target domain. Finally, we conduct comprehensive experiments on seven public datasets to evaluate and analyze the effectiveness of SAMGPT.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。