解决多模态学习中主导模态压制其他模态的问题
Modality-Aware SAM: Sharpness-Aware-Minimization Driven Gradient Modulation for Harmonized Multimodal Learning
- 用Shapley值识别主导模态,动态调整学习重点
- 通过梯度调制增强模型对主导模态的鲁棒性
- 适用于早融合与晚融合,提升整体性能平衡
在多模态学习中,主导模态常压制其他模态,限制泛化能力。本文提出一种模型无关的框架M-SAM,适用于多种模态及早融合、晚融合场景。每轮迭代中,M-SAM分三步优化:首先基于Shapley值评估各模态对准确率的贡献,识别主导模态;其次分解损失曲面,通过梯度调制强化模型对主导模态的鲁棒性;最后使用调制后的梯度进行反向传播更新权重。该机制在保障主导模态性能的同时,促进其他模态贡献,使模型能有效挖掘互补特征,提升整体表现。在四个不同数据集上的大量实验表明,M-SAM优于当前最先进的优化与梯度调控方法,显著提升了多模态学习的平衡性与性能。
原文摘要 · Abstract (English)
In multimodal learning, dominant modalities often overshadow others, limiting generalization. We propose Modality-Aware Sharpness-Aware Minimization (M-SAM), a model-agnostic framework that applies to many modalities and supports early and late fusion scenarios. In every iteration, M-SAM in three steps optimizes learning. \textbf{First, it identifies the dominant modality} based on modalities' contribution in the accuracy using Shapley. \textbf{Second, it decomposes the loss landscape}, or in another language, it modulates the loss to prioritize the robustness of the model in favor of the dominant modality, and \textbf{third, M-SAM updates the weights} by backpropagation of modulated gradients. This ensures robust learning for the dominant modality while enhancing contributions from others, allowing the model to explore and exploit complementary features that strengthen overall performance. Extensive experiments on four diverse datasets show that M-SAM outperforms the latest state-of-the-art optimization and gradient manipulation methods and significantly balances and improves multimodal learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。