arXiv:2602.09507cs.LG2026-02被引 4

解决多模态表示学习中的对齐与均匀性冲突,提升跨模态一致性。

Towards Uniformity and Alignment for Multimodal Representation Learning

  • 解耦对齐与均匀性,缓解多模态间分布差距
  • 理论证明方法可有效缩小多模态分布差异
  • 在检索与生成任务中均取得稳定提升

多模态表示学习旨在构建一个共享嵌入空间,使异构模态语义对齐。尽管表现优异,基于InfoNCE的目标函数引入了固有冲突,导致模态间分布差异。本文识别出两种在模态数量增加时加剧的冲突:(i) 对齐-均匀性冲突,即均匀性排斥削弱成对对齐;(ii) 内部对齐冲突,多个模态对齐产生相互竞争的方向。为此,我们提出一种原则性的对齐与均匀性解耦方法,提供无需任务特定模块的无冲突多模态学习方案,支持判别与生成双重场景。我们进一步提供理论保证,证明该方法是多模态分布间全局Hölder散度的有效代理,从而减小模态间分布差距。在检索与UnCLIP式生成任务上的大量实验表明方法具有一致性增益。

原文摘要 · Abstract (English)

Multimodal representation learning aims to construct a shared embedding space in which heterogeneous modalities are semantically aligned. Despite strong empirical results, InfoNCE-based objectives introduce inherent conflicts that yield distribution gaps across modalities. In this work, we identify two conflicts in the multimodal regime, both exacerbated as the number of modalities increases: (i) an alignment-uniformity conflict, whereby the repulsion of uniformity undermines pairwise alignment, and (ii) an intra-alignment conflict, where aligning multiple modalities induces competing alignment directions. To address these issues, we propose a principled decoupling of alignment and uniformity for multimodal representations, providing a conflict-free recipe for multimodal learning that simultaneously supports discriminative and generative use cases without task-specific modules. We then provide a theoretical guarantee that our method acts as an efficient proxy for a global Hölder divergence over multiple modality distributions, and thus reduces the distribution gap among modalities. Extensive experiments on retrieval and UnCLIP-style generation demonstrate consistent gains.

多模态表示学习对齐生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。