arXiv:2606.11614cs.LGcs.AI2026-06中稿 · CVPR

提出新方法动态分解多模态交互,提升模型理解能力。

Information-Theoretic Decomposition for Multimodal Interaction Learning

论文配图:Information-Theoretic Decomposition for Multimodal Interaction Learning
图 1 · 摘自论文原文
  • 用变分分解架构分离多模态中的冗余、独特和协同信息
  • 在多个任务上显著优于传统方法,性能持续领先
  • 适合需要精细理解跨模态关系的场景,如医疗诊断

多模态学习的核心在于捕捉不同模态间的冗余、独特和协同信息,这些共同构成多模态交互。一个关键但未被充分研究的挑战是:这些隐含交互在样本间具有动态变化特性。本文首次系统性地进行信息论分析,揭示为何学习样本级动态交互对有效多模态学习至关重要。分析还指出,传统范式存在缺陷:模态集成方法难以捕捉协同作用,而联合学习常低估冗余信息。为此,我们提出基于分解的多模态交互学习(DMIL),一种新型范式,可显式建模并学习样本特定的交互。首先设计变分分解架构以分离交互成分;其次采用新学习策略,在微调过程中利用这些显式交互组件,实现全面交互学习。在多种任务与架构上的大量实验表明,DMIL通过适应整体样本级交互,持续取得更优性能。该框架灵活且普适,确立了以交互为中心的多模态学习新范式。代码已公开于 https://github.com/GeWu-Lab/DMIL。

原文摘要 · Abstract (English)

Multimodal learning hinges on capturing redundant, unique, and synergistic information across modalities, which collectively constitute multimodal interactions. A critical yet underexplored challenge is that these implicit interactions vary dynamically across samples. In this work, we present the first systematic, information-theoretic analysis highlighting why learning these dynamic, sample-specific interactions is critical for effective multimodal learning. Our analysis further reveals deficits in conventional paradigms at learning these distinct interaction types: modality ensemble approaches struggle to capture synergy, while joint learning paradigms often under-utilize redundant information. This highlights the need for an approach that can adaptively learn from different interaction types on a per-sample basis. To this end, we propose Decomposition-based Multimodal Interaction Learning (DMIL), a novel paradigm that explicitly models and learns from sample-specific interactions. First, we design a variational decomposition architecture to isolate the constituent interaction components. Second, we employ a new learning strategy that leverages these explicit interaction components in a fine-tuning process to achieve comprehensive interaction learning. Extensive experiments across diverse tasks and architectures demonstrate that DMIL consistently achieves superior performance by adapting to holistic sample-specific interactions. Our framework is flexible and broadly applicable, establishing an interaction-centric paradigm for multimodal learning. The code is available at https://github.com/GeWu-Lab/DMIL.

多模态学习信息论交互建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。