arXiv:2503.13834cs.CV2025-03NAACL被引 13

解决视觉语言模型偏倚问题,让双模态更均衡协作。

See-Saw Modality Balance: See Gradient, and Sew Impaired Vision-Language Balance to Mitigate Dominant Modality Bias

  • 通过梯度重加权和方向对齐,缓解单模态主导现象。
  • 在多个数据集上显著降低对单一模态的依赖性。
  • 适合关注多模态公平性与鲁棒性的研究者使用。

视觉语言(VL)模型在各类任务中表现强劲,但常过度依赖某一模态,导致‘主导模态偏倚’,尤其在某模态受损时性能大幅下降。本研究分析该偏倚下的模型行为,理论证明梯度不对齐或幅值差异会阻碍损失的平衡收敛。为此提出新框架BalGrad,包含跨模态梯度重加权、基于模态贡献调整KL散度梯度,以及跨任务梯度投影以对齐任务方向。在UPMC Food-101、Hateful Memes和MM-IMDb数据集上的实验表明,BalGrad能有效缓解对特定模态的过度依赖。

原文摘要 · Abstract (English)

Vision-language (VL) models have demonstrated strong performance across various tasks. However, these models often rely on a specific modality for predictions, leading to "dominant modality bias.'' This bias significantly hurts performance, especially when one modality is impaired. In this study, we analyze model behavior under dominant modality bias and theoretically show that unaligned gradients or differences in gradient magnitudes prevent balanced convergence of the loss. Based on these findings, we propose a novel framework, BalGrad to mitigate dominant modality bias. Our approach includes inter-modality gradient reweighting, adjusting the gradient of KL divergence based on each modality's contribution, and inter-task gradient projection to align task directions in a non-conflicting manner. Experiments on UPMC Food-101, Hateful Memes, and MM-IMDb datasets confirm that BalGrad effectively alleviates over-reliance on specific modalities when making predictions.

多模态偏倚缓解视觉语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。