arXiv:2505.19996cs.LG2025-05ICML被引 13

提出OMIB框架,动态调节多模态信息瓶颈的正则化权重

Learning Optimal Multimodal Information Bottleneck Representations

  • 基于理论推导设定正则化权重边界,确保最优信息瓶颈可达成
  • 按模态动态调整权重,缓解任务相关信息不平衡问题
  • 在合成数据和下游任务中验证理论性质与性能优势

利用高质量的多模态联合表示可显著提升多种机器学习应用的模型性能。近期基于多模态信息瓶颈(MIB)原则的方法通过正则化生成最优MIB,保留最大任务相关信息并去除冗余信息。然而,这些方法常采用经验设定的正则化权重,并忽视模态间任务相关性分布不均的问题,限制了其达到最优MIB的能力。为此,本文提出一种新型多模态学习框架——最优多模态信息瓶颈(OMIB),其优化目标通过理论推导的权重范围保证最优MIB的可达性。OMIB进一步通过按模态动态调整正则化权重,促进所有任务相关信号的充分保留。此外,本文建立了坚实的信息论基础,并在变分近似框架下实现高效计算。最后,我们在合成数据上验证了OMIB的理论性质,并在多个下游任务中展示了其优于现有基准方法的性能。

原文摘要 · Abstract (English)

Leveraging high-quality joint representations from multimodal data can greatly enhance model performance in various machine-learning based applications. Recent multimodal learning methods, based on the multimodal information bottleneck (MIB) principle, aim to generate optimal MIB with maximal task-relevant information and minimal superfluous information via regularization. However, these methods often set ad hoc regularization weights and overlook imbalanced task-relevant information across modalities, limiting their ability to achieve optimal MIB. To address this gap, we propose a novel multimodal learning framework, Optimal Multimodal Information Bottleneck (OMIB), whose optimization objective guarantees the achievability of optimal MIB by setting the regularization weight within a theoretically derived bound. OMIB further addresses imbalanced task-relevant information by dynamically adjusting regularization weights per modality, promoting the inclusion of all task-relevant information. Moreover, we establish a solid information-theoretical foundation for OMIB's optimization and implement it under the variational approximation framework for computational efficiency. Finally, we empirically validate the OMIB's theoretical properties on synthetic data and demonstrate its superiority over the state-of-the-art benchmark methods in various downstream tasks.

多模态学习信息瓶颈动态正则化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。