用几何体积正则化提升多模态对齐效果,让文本视频音频更精准匹配。
MOVER: Multimodal Optimal Transport with Volume-based Embedding Regularization
- 基于最优传输软对齐+体积最小化,实现跨模态一致对齐。
- 在零样本和微调设置下均超越现有最佳方法。
- 适合需要强泛化能力的多模态检索任务。
多模态学习近年主要依赖成对对比目标,将文本、视频、音频等模态对齐到共享嵌入空间。这类方法在双模态场景有效,但难以推广到多模态,且高维空间中缺乏语义结构。本文提出MOVER框架,结合基于最优传输的软对齐与基于体积的几何正则化,构建语义对齐且结构化的多模态表示。通过传输引导的匹配机制与几何体积最小化目标(GAVE),MOVER实现模态无关的一致对齐。在文本-视频-音频检索任务上的实验表明,MOVER在零样本和微调设置下显著优于现有最先进方法。额外分析显示其对未见模态组合具有更强泛化能力,且嵌入空间结构一致性更强。
原文摘要 · Abstract (English)
Recent advances in multimodal learning have largely relied on pairwise contrastive objectives to align different modalities, such as text, video, and audio, in a shared embedding space. While effective in bi-modal setups, these approaches struggle to generalize across multiple modalities and often lack semantic structure in high-dimensional spaces. In this paper, we propose MOVER, a novel framework that combines optimal transport-based soft alignment with volume-based geometric regularization to build semantically aligned and structured multimodal representations. By integrating a transport-guided matching mechanism with a geometric volume minimization objective (GAVE), MOVER encourages consistent alignment across all modalities in a modality-agnostic manner. Experiments on text-video-audio retrieval tasks demonstrate that MOVER significantly outperforms prior state-of-the-art methods in both zero-shot and finetuned settings. Additional analysis shows improved generalization to unseen modality combinations and stronger structural consistency in the learned embedding space.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。