arXiv:2606.12867cs.LG2026-06

通过频域分解分离图结构与模态语义,提升多模态图学习效果。

SMGFM: Spectral Multimodal Graph Pretraining for Multimodal-Attributed Graphs

论文配图:SMGFM: Spectral Multimodal Graph Pretraining for Multimodal-Attributed Graphs
图 1 · 摘自论文原文
  • 基于图频域特性,将不同模态信号分解为高低频成分。
  • 在频段级别分配语义角色,避免跨模态过度对齐。
  • 适合多模态图数据建模,尤其文本图像融合场景。

多模态属性图(MAGs)将图结构与文本、图像等多模态语义结合。传统图学习通过拓扑与节点特征耦合来理解语义,但在MAG中,结构诱导语义与模态内生语义的作用机制不同:前者促进关系一致性,后者保留局部细粒度差异,不应统一平滑或对齐。因此,关键挑战在于跨模态融合前识别语义角色。本文提出SMGFM,一种基于频域特性的多模态图预训练框架,利用图频率变化作为先验——低频成分捕捉拓扑一致语义,高频成分保留模态特异性语义。SMGFM通过可扩展的切比雪夫滤波器构建频段分辨的模态令牌,基于拓扑条件路由估计耦合可靠性,并在频段-模态间交互后再融合。其频段路由目标既对齐平滑共识路径,又保留模态特异路径,缓解空间域混叠与统一对齐问题。在多个MAG数据集上的实验表明,SMGFM在图级和模态级任务上均达到当前最优性能。

原文摘要 · Abstract (English)

Multimodal-attributed graphs (MAGs) couple graph topology with node semantics from text, images, and other modalities. Traditional graph learning contextualizes node semantics by coupling topology with node features. However, this coupling design becomes troublesome in MAGs, where structure-induced and modality-intrinsic semantics may contribute differently to downstream tasks. Structure-induced semantics promote relational consistency through smooth topological variation, whereas modality-intrinsic semantics often encode local, fine-grained distinctions that should not be uniformly smoothed or aligned. Therefore, the key challenge is to identify semantic roles before cross-modal fusion. To this end, we leverage graph-frequency variation as a prior, where low-frequency components capture topology-consistent semantics and high-frequency components preserve modality-specific semantics. Based on this intuition, we propose SMGFM, a spectral multimodal graph pretraining framework that decomposes each modality-specific node signal into graph-frequency bands and assigns band-level semantic roles before cross-modal interaction. Concretely, SMGFM constructs frequency-resolved modality tokens with scalable Chebyshev filters, estimates their coupling reliability through topology-conditioned routing, and performs band-modality interaction before fusion. Its frequency-routed objectives align smooth consensus routes while preserving modality-specific routes, mitigating spatial-domain entanglement and uniform cross-modal alignment. Extensive experiments conducted on the MAG datasets demonstrate that SMGFM achieves state-of-the-art performance across graph-level and modality-level tasks.

多模态图频域分析预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。