发现多模态模型在决策层仍存在偏见,提出应动态分配权重以平衡模态贡献。
Revisit Modality Imbalance at the Decision Layer
- 在融合阶段引入自适应权重分配机制,缓解决策层的模态偏见。
- 实验显示即使预训练充分,音频模态仍主导决策结果(如CREMAD数据集)。
- 适合关注多模态公平性与模型可解释性的研究者阅读。
多模态学习通过融合不同模态信息提升性能,但常面临模态不平衡问题,即主导模态在联合优化中压制弱模态。本文揭示,这种不平衡不仅存在于表征学习阶段,也显著体现在决策层。在音频-视觉数据集CREMAD和Kinetic-Sounds上的实验表明,即使经过充分预训练和均衡优化,模型仍系统性偏向特定模态(如音频)。进一步分析显示,该偏见源于特征空间与决策权重分布的内在差异,而非仅由优化动态导致。我们指出,在融合阶段使用未经校准的模态输出会导致决策层权重失衡,阻碍弱模态有效贡献。为此,建议未来多模态系统应在决策层引入自适应权重分配机制,依据各模态实际能力实现相对平衡。
原文摘要 · Abstract (English)
Multimodal learning integrates information from different modalities to enhance model performance, yet it often suffers from modality imbalance, where dominant modalities overshadow weaker ones during joint optimization. This paper reveals that such an imbalance not only occurs during representation learning but also manifests significantly at the decision layer. Experiments on audio-visual datasets (CREMAD and Kinetic-Sounds) show that even after extensive pretraining and balanced optimization, models still exhibit systematic bias toward certain modalities, such as audio. Further analysis demonstrates that this bias originates from intrinsic disparities in feature-space and decision-weight distributions rather than from optimization dynamics alone. We argue that aggregating uncalibrated modality outputs at the fusion stage leads to biased decision-layer weighting, hindering weaker modalities from contributing effectively. To address this, we propose that future multimodal systems should focus more on incorporate adaptive weight allocation mechanisms at the decision layer, enabling relative balanced according to the capabilities of each modality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。