arXiv:2602.04116cs.LGcs.AI2026-02被引 3

提出PLANET框架,解决多模态图模型中模态交互与对齐难题。

Toward Effective Multimodal Graph Foundation Model: A Divide-and-Conquer Based Approach

  • 分治策略解耦模态交互与对齐,提升跨模态理解能力
  • 在多个图任务上优于现有基线,性能显著提升
  • 适合需要融合文本、图像等多模态信息的图学习场景

图基础模型(GFMs)在跨领域泛化方面取得显著进展,但主要聚焦于文本属性图(TAGs),而多模态属性图(MAGs)尚未被充分挖掘。构建多模态图基础模型(MGFMs)可利用MAG中丰富的多模态信息,扩展下游任务适用范围。尽管近期研究整合了多种模态信息,我们的实证分析揭示现有MGFMs存在两大根本缺陷:(1)未显式建模模态交互,难以捕捉超越简单聚合的复杂跨模态语义;(2)模态对齐表现不佳,无法有效弥合不同模态空间间的显著语义差异。为此,我们提出PLANET(图拓扑感知模态交互与对齐框架),采用分治策略,在不同粒度上解耦模态交互与对齐。在嵌入粒度,通过嵌入级域门控(EDG)自适应注入拓扑感知的跨模态上下文,实现局部语义增强;在节点粒度,通过节点级离散化检索(NDR)构建离散语义表示空间(DSRS),实现全局模态对齐。大量实验表明,PLANET在多种图中心任务和多模态生成任务中显著优于当前最优基线。

原文摘要 · Abstract (English)

Graph Foundation Models (GFMs) have achieved remarkable success in generalizing across diverse domains. However, they mainly focus on Text-Attributed Graphs (TAGs), leaving Multimodal-Attributed Graphs (MAGs) largely untapped. Developing Multimodal Graph Foundation Models (MGFMs) allows for leveraging the rich multimodal information in MAGs, and extends applicability to broader types of downstream tasks. While recent MGFMs integrate diverse modality information, our empirical investigation reveals two fundamental limitations of existing MGFMs: (1)they fail to explicitly model modality interaction, essential for capturing intricate cross-modal semantics beyond simple aggregation, and (2)they exhibit sub-optimal modality alignment, which is critical for bridging the significant semantic disparity between distinct modal spaces. To address these challenges, we propose PLANET (graPh topoLogy-aware modAlity iNteraction and alignmEnT), a novel framework employing a Divide-and-Conquer strategy to decouple modality interaction and alignment across distinct granularities. At the embedding granularity, (1)Embedding-wise Domain Gating (EDG) performs local semantic enrichment by adaptively infusing topology-aware cross-modal context, achieving modality interaction. At the node granularity, (2)Node-wise Discretization Retrieval (NDR) ensures global modality alignment by constructing a Discretized Semantic Representation Space (DSRS) to bridge modality gaps. Extensive experiments demonstrate that PLANET significantly outperforms state-of-the-art baselines across diverse graph-centric and multimodal generative tasks.

多模态图图神经网络模态对齐分治策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。