提出不对称增强方法,解决多模态学习中主导模态压制弱模态的问题。
Asymmetric Reinforcing against Multi-modal Representation Bias
- 动态增强弱模态,同时保留强模态表征能力
- 通过条件互信息优化,减少信息损失
- 适合处理模态贡献动态变化的现实场景
多模态学习的优势在于整合多源信息,提供丰富全面的洞察。然而在真实场景中,多模态系统常面临模态贡献动态变化的问题,不同模态的主导性随环境变化,导致多模态学习性能下降。现有方法主要通过增强弱模态来平衡表示偏差,但往往从部分模态视角优化,易导致强模态性能下降。为此,我们提出对抗多模态表示偏差的不对称增强方法(ARM)。ARM通过条件互信息动态强化弱模态,同时保持对强模态的表征能力。我们进一步分析发现,仅优化特定模态会导致信息丢失,阻碍充分利用多模态数据优势。通过探索模态主导性并缩小模态间贡献差距,显著提升了多模态学习性能,在缓解不平衡多模态学习方面取得明显进展。
原文摘要 · Abstract (English)
The strength of multimodal learning lies in its ability to integrate information from various sources, providing rich and comprehensive insights. However, in real-world scenarios, multi-modal systems often face the challenge of dynamic modality contributions, the dominance of different modalities may change with the environments, leading to suboptimal performance in multimodal learning. Current methods mainly enhance weak modalities to balance multimodal representation bias, which inevitably optimizes from a partialmodality perspective, easily leading to performance descending for dominant modalities. To address this problem, we propose an Asymmetric Reinforcing method against Multimodal representation bias (ARM). Our ARM dynamically reinforces the weak modalities while maintaining the ability to represent dominant modalities through conditional mutual information. Moreover, we provide an in-depth analysis that optimizing certain modalities could cause information loss and prevent leveraging the full advantages of multimodal data. By exploring the dominance and narrowing the contribution gaps between modalities, we have significantly improved the performance of multimodal learning, making notable progress in mitigating imbalanced multimodal learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。