用图信号平滑思想,让冻结的图文嵌入更准地对齐。
GOMA: Toward Structure-Driven Multimodal Alignment from a Graph Signal Smoothing Perspective

- 将图文嵌入看作图信号,分步设计传播路径与平滑深度
- 在7个基准上达到顶尖或持平最佳效果,稳定性显著提升
- 适合需要提升图文对齐精度的检索任务使用
多模态对齐通常通过CLIP式双编码器从孤立的图文对偶中学习,忽略了实体间的关联上下文。多模态属性图(MAG)中,节点携带多模态属性,边编码语料结构,为优化冻结的视觉-语言嵌入提供了自然场景。该优化面临挑战:视觉、文本和跨模态关系常引发不同邻域几何,而无限制的图传播易导致表示过平滑。有效利用图结构需同时打破模态特异性拓扑障碍、控制平滑程度,并在语义边界坍塌前保留有用信息。我们提出图优化多模态对齐(GOMA),一种结构驱动的后对齐框架,将冻结的多模态嵌入视为图信号,通过统一的检索导向设计满足上述需求。GOMA解耦三个关键设计:消息流向何处、多模态证据如何传播、应保留何种平滑深度。具体而言,它学习模态感知的传播算子,执行无对角线跨模态捷径的有限步耦合平滑,并自适应读出节点特定的平滑轨迹以保留有益平滑。所有实验遵循归纳式MAG检索协议,图仅作为未标注上下文,且移除对角自配对边。在七个MAG基准上,GOMA实现最优或并列最优的检索性能,且显著优于最强的图对比方法,证明了MAG结构可作为冻结多模态嵌入的有效后编码器。
原文摘要 · Abstract (English)
Multimodal alignment is commonly learned from isolated image-text pairs via CLIP-style dual encoders, leaving the relational context among entities largely unused. Multimodal attributed graphs (MAGs), where nodes carry multimodal attributes and edges encode corpus structure, provide a natural setting for refining frozen vision-language embeddings. This refinement is challenging: visual, textual, and cross-modal relations often induce different neighborhood geometries, while unrestricted graph propagation can quickly over-smooth retrieval representations. Effectively leveraging graph context therefore requires simultaneously breaking modality-specific topological barriers, controlling the smoothing regime, and preserving informative smoothing before semantic boundaries collapse. We propose Graph-Optimized Multimodal Alignment (GOMA), a structure-driven post-alignment framework that views frozen multimodal embeddings as graph signals and addresses these requirements through a unified retrieval-oriented design. GOMA decouples three key design choices: where messages should flow, how multimodal evidence should propagate, and which smoothing depth should be retained. Concretely, it learns modality-aware propagation operators, performs finite-step coupled smoothing without diagonal cross-modal shortcuts, and adaptively reads out node-specific smoothing trajectories to preserve useful smoothing before collapse. All experiments follow a transductive MAG retrieval protocol where the graph serves only as unlabeled context and diagonal self-pair edges are removed. On seven MAG benchmarks, GOMA achieves state-of-the-art or tied state-of-the-art retrieval and remains substantially more stable than the strongest graph competitor, demonstrating that MAG structure can serve as an effective post-encoder for frozen multimodal embeddings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。