arXiv:2606.09853cs.LGcs.IT2026-06

提出新训练目标,让多模态模型更专注捕捉跨模态协同信息。

SynIB: Informational Bottleneck for Maximizing Synergy in Multimodal Learning

论文配图:SynIB: Informational Bottleneck for Maximizing Synergy in Multimodal Learning
图 1 · 摘自论文原文
  • 基于信息瓶颈理论设计目标函数,直接优化跨模态协同
  • 在合成任务和真实数据集上提升协同相关样本准确率最高7.8%
  • 适合需要深度跨模态理解的任务,如情感分析与讽刺识别

多模态学习的核心目标是捕捉协同效应:仅在联合使用多种模态时才出现的、单个模态无法提供的任务相关信息。现有方法主要通过构建更复杂的融合架构来实现,我们提出一个互补方向:改变训练目标本身。标准训练往往强调单模态或冗余信息,难以学习跨模态推理。本文通过信息论形式化多模态协同,并提出可扩展的协同信息瓶颈(SynIB)目标。为强化协同学习,SynIB要求模型在所有模态下均能准确预测,同时在遮蔽任一模态时惩罚过高置信度——这表明其依赖单模态线索而非跨模态交互。在两个场景中验证:在构造已知协同效应的合成XOR任务中,标准训练失败,而SynIB成功恢复;在五个真实世界基准上,包括三个MultiBench情感任务、带有CLIP-ViT和DeBERTa主干的Hateful Memes,以及我们引入的可控讽刺扩展CREMA-D,SynIB在协同依赖样本上最高提升7.8%准确率,整体准确率最高提升3.8%。

原文摘要 · Abstract (English)

A central objective in multimodal learning is to capture synergy: task-relevant information that arises only from the joint use of multiple modalities, and is not available from any single modality alone. While most approaches operate at the architectural level through larger or more complex fusion models, we propose a complementary axis: shaping the training objective itself. Standard training often emphasizes unimodal or redundant information, falling short on examples that require cross-modal reasoning. We formalize multimodal synergy through information theory and introduce the Synergistic Information Bottleneck (SynIB), a scalable objective that targets synergy directly. To prioritize learning synergy, SynIB motivates the model to predict accurately from all modalities while penalizing confidence when information from any modality is withheld. Alongside the standard task loss, the model runs forward passes with one modality masked at a time and is penalized for remaining confident, which would indicate reliance on unimodal cues rather than cross-modal interactions. We validate SynIB in two regimes. On synthetic XOR tasks where the ground-truth synergy is known by construction, standard training fails to recover it while SynIB does. On five real-world benchmarks, including three MultiBench affective tasks, Hateful Memes with CLIP-ViT and DeBERTa backbones, and a controllable irony extension of CREMA-D we introduce, SynIB improves accuracy on synergy-dependent examples by up to 7.8% and overall accuracy by up to 3.8%.

多模态学习协同效应信息瓶颈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。