arXiv:2503.11892cs.CV2025-03中稿 · ICLR被引 33

提出分层对齐框架,分离多模态特征中的共性与个性信息。

DecAlign: Hierarchical Cross-Modal Alignment for Decoupled Multimodal Representation Learning

  • 通过原型引导的最优传输策略处理模态异构性
  • 在四个基准上五项指标均优于现有方法
  • 适合需要保留模态特异性的多模态任务

多模态表征学习旨在捕捉跨多种模态的共享与互补语义信息。然而,不同模态固有的异质性给有效的跨模态协作与融合带来挑战。为此,我们提出DecAlign,一种新型分层跨模态对齐框架,将多模态表征解耦为模态独有(异质)和模态共有(同质)特征。针对异质性,采用基于高斯混合建模和多边际传输计划的原型引导最优传输对齐策略,缓解分布差异的同时保留模态独有特性。为增强同质性,通过隐空间分布匹配与最大均值差异正则化确保跨模态语义一致性。此外,引入多模态Transformer提升高层语义特征融合,进一步降低跨模态不一致性。在四个广泛使用的多模态基准上的大量实验表明,DecAlign在五项指标上持续优于现有最先进方法。结果凸显其在提升跨模态对齐与语义一致性的同时保持模态独有特征的有效性,标志着多模态表征学习的重要进展。项目主页:https://taco-group.github.io/DecAlign。

原文摘要 · Abstract (English)

Multimodal representation learning aims to capture both shared and complementary semantic information across multiple modalities. However, the intrinsic heterogeneity of diverse modalities presents substantial challenges to achieve effective cross-modal collaboration and integration. To address this, we introduce DecAlign, a novel hierarchical cross-modal alignment framework designed to decouple multimodal representations into modality-unique (heterogeneous) and modality-common (homogeneous) features. For handling heterogeneity, we employ a prototype-guided optimal transport alignment strategy leveraging gaussian mixture modeling and multi-marginal transport plans, thus mitigating distribution discrepancies while preserving modality-unique characteristics. To reinforce homogeneity, we ensure semantic consistency across modalities by aligning latent distribution matching with Maximum Mean Discrepancy regularization. Furthermore, we incorporate a multimodal transformer to enhance high-level semantic feature fusion, thereby further reducing cross-modal inconsistencies. Our extensive experiments on four widely used multimodal benchmarks demonstrate that DecAlign consistently outperforms existing state-of-the-art methods across five metrics. These results highlight the efficacy of DecAlign in enhancing superior cross-modal alignment and semantic consistency while preserving modality-unique features, marking a significant advancement in multimodal representation learning scenarios. Our project page is at https://taco-group.github.io/DecAlign.

多模态学习表征解耦跨模态对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。