arXiv:2602.14983cs.LG2026-02

提出COrAL框架,显式分离多模态中的共享、特有和协同信息。

Orthogonalized Multimodal Contrastive Learning with Asymmetric Masking for Structured Representations

  • 双路径结构加正交约束,解耦跨模态共享与模态特有特征。
  • 采用非对称掩码促进跨模态依赖学习,增强协同信息建模。
  • 在多个数据集上表现稳定且优于主流方法,适合追求鲁棒多模态表示的场景。

多模态学习旨在整合异构来源的信息,这些信息可能在模态间共享、仅属于特定模态,或仅通过交互产生。尽管自监督多模态对比学习已取得显著进展,现有方法大多聚焦冗余的跨模态信号,常忽略模态特有(独特)和交互驱动(协同)信息。近期改进虽拓展了视角,但或未能显式建模协同交互,或以纠缠方式学习不同信息成分,导致表征不完整并引发信息泄露。我们提出 extbf{COrAL},一种原理性框架,能显式且同时保留冗余、独特和协同信息。COrAL采用双路径架构结合正交性约束,解耦共享与模态特有特征,确保信息成分清晰分离。为强化协同建模,引入具有互补视图特性的非对称掩码,迫使模型推断跨模态依赖,而非仅依赖冗余线索。在合成基准与多样化的MultiBench数据集上的大量实验表明,COrAL始终达到或超越当前最优方法,且运行间性能方差极低。结果表明,显式建模多模态信息全谱可获得更稳定、可靠且全面的嵌入表示。

原文摘要 · Abstract (English)

Multimodal learning seeks to integrate information from heterogeneous sources, where signals may be shared across modalities, specific to individual modalities, or emerge only through their interaction. While self-supervised multimodal contrastive learning has achieved remarkable progress, most existing methods predominantly capture redundant cross-modal signals, often neglecting modality-specific (unique) and interaction-driven (synergistic) information. Recent extensions broaden this perspective, yet they either fail to explicitly model synergistic interactions or learn different information components in an entangled manner, leading to incomplete representations and potential information leakage. We introduce \textbf{COrAL}, a principled framework that explicitly and simultaneously preserves redundant, unique, and synergistic information within multimodal representations. COrAL employs a dual-path architecture with orthogonality constraints to disentangle shared and modality-specific features, ensuring a clean separation of information components. To promote synergy modeling, we introduce asymmetric masking with complementary view-specific patterns, compelling the model to infer cross-modal dependencies rather than rely solely on redundant cues. Extensive experiments on synthetic benchmarks and diverse MultiBench datasets demonstrate that COrAL consistently matches or outperforms state-of-the-art methods while exhibiting low performance variance across runs. These results indicate that explicitly modeling the full spectrum of multimodal information yields more stable, reliable, and comprehensive embeddings.

多模态对比学习表征解耦协同建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。