arXiv:2601.22856cs.LG2026-01

用最优传输对齐多模态图的结构与语义,提升表示学习效果

OptiMAG: Structure-Semantic Alignment via Unbalanced Optimal Transport

  • 基于非平衡最优传输,用融合格罗莫夫-沃瑟斯坦距离对齐跨模态结构
  • 在节点分类、链接预测等任务上超越基线模型,生成任务性能显著提升
  • 可无缝嵌入现有模型,适合多模态图学习研究者使用

多模态属性图(MAGs)通过在节点上整合文本、图像等多模态信息,广泛用于建模复杂系统。然而我们发现,不同模态嵌入隐含的语义结构与显式的图结构之间存在不一致:例如,某些邻居在一种模态中相近,在另一模态中却相远。由于现有方法通常在固定的显式图结构上进行消息传递,会无意中聚合不相似特征,引入模态特异性噪声,阻碍有效节点表示学习。为此,我们提出OptiMAG,一种基于非平衡最优传输的正则化框架。OptiMAG采用融合格罗莫夫-沃瑟斯坦距离,显式引导局部邻域内跨模态的结构一致性,有效缓解结构-语义冲突。此外,通过KL散度惩罚实现对跨模态不一致性的自适应处理。该框架可无缝集成至现有多模态图模型中,作为有效的即插即用正则项。实验表明,OptiMAG在多个任务上持续优于基线模型,涵盖以图为中心的任务(如节点分类、链接预测)和以多模态为中心的生成任务(如graph2text、graph2image)。源代码将在录用后公开。

原文摘要 · Abstract (English)

Multimodal Attributed Graphs (MAGs) have been widely adopted for modeling complex systems by integrating multi-modal information, such as text and images, on nodes. However, we identify a discrepancy between the implicit semantic structure induced by different modality embeddings and the explicit graph structure. For instance, neighbors in the explicit graph structure may be close in one modality but distant in another. Since existing methods typically perform message passing over the fixed explicit graph structure, they inadvertently aggregate dissimilar features, introducing modality-specific noise and impeding effective node representation learning. To address this, we propose OptiMAG, an Unbalanced Optimal Transport-based regularization framework. OptiMAG employs the Fused Gromov-Wasserstein distance to explicitly guide cross-modal structural consistency within local neighborhoods, effectively mitigating structural-semantic conflicts. Moreover, a KL divergence penalty enables adaptive handling of cross-modal inconsistencies. This framework can be seamlessly integrated into existing multimodal graph models, acting as an effective drop-in regularizer. Experiments demonstrate that OptiMAG consistently outperforms baselines across multiple tasks, ranging from graph-centric tasks (e.g., node classification, link prediction) to multimodal-centric generation tasks (e.g., graph2text, graph2image). The source code will be available upon acceptance.

多模态图最优传输结构对齐表示学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。