arXiv:2606.14172cs.LGcs.CV2026-06

提出新模型CoMAG,让多模态图在不同任务中自适应保留关键信息。

Context-aware Modality-Topology Co-Alignment for Multimodal Attributed Graphs

论文配图:Context-aware Modality-Topology Co-Alignment for Multimodal Attributed Graphs
图 1 · 摘自论文原文
  • 根据语义一致性动态构建可靠图上下文,避免固定拓扑干扰
  • 通过多跳令牌对齐保持各模态特征,提升跨模态匹配精度
  • 适合需要精细跨模态对齐的图文生成与结构预测任务

多模态属性图(MAGs)通过融合图结构与异构属性(如文本、图像)建模真实世界实体,支持图级任务和模态级任务。现有方法常依赖固定图上下文或统一特征融合,导致传播方式通用化、模态信息压缩过度,难以满足多样化任务需求。为此,我们提出CoMAG,一个统一的MAG主干网络,可学习任务自适应的可靠上下文并实现模态保持对齐。CoMAG首先通过多模态语义一致性估计边可靠性,补充原始拓扑并筛选上下文组件;随后进行模态保持的多跳令牌对齐,维护各模态的多跳路径,跨模态匹配令牌并解耦共享与私有表示。该方法在一次前向传播中生成图与模态表示,同时保留模态特异性线索。我们分析了稳定传播、过平滑抑制与模态坍缩控制。在九个OpenMAG数据集上,相比特征仅用、图仅用、多模态及统一MAG基线,CoMAG在图级预测、模态匹配和图条件生成任务中均取得最优表现,验证了任务自适应可靠上下文与模态保持对齐的有效性,且维持稀疏边-线性复杂度。

原文摘要 · Abstract (English)

Multimodal Attributed Graphs (MAGs) model real-world entities by coupling graph topology with heterogeneous attributes such as text and images. They support graph-centric tasks requiring structural and class-discriminative representations, and modality-centric tasks requiring fine-grained cross-modal correspondence. However, existing MAG methods often rely on fixed graph contexts or uniformly fused representations, causing task-agnostic propagation and over-compressed fusion that hinder diverse task requirements and modality-specific evidence preservation. To address this, we propose CoMAG, a unified MAG backbone that learns task-adaptive reliable contexts and modality-preserving alignment within them. CoMAG first conducts Reliable Context Learning by estimating edge reliability from multimodal semantic consistency, complementing raw topology with semantic neighbors, and selecting context components through a task-aware gate. It then performs Modality-preserving Hop-token Alignment by maintaining modality-specific multi-hop trajectories, matching modality-hop tokens across modalities, and decoupling shared and private representations. Thus, CoMAG produces graph and modality representations from one forward pass while retaining modality-specific cues. We further analyze stable propagation, over-smoothing mitigation, and modality-collapse control. Experiments on nine OpenMAG datasets compare CoMAG with feature-only, graph-only, multimodal, and unified MAG baselines across graph-level prediction, modality matching, and graph-conditioned generation. Results show that CoMAG achieves the best reported performance, demonstrating that task-adaptive reliable contexts and modality-preserving alignment improve structural prediction, cross-modal matching, and graph-conditioned generation while retaining sparse edge-linear complexity.

多模态图跨模态对齐图神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。