arXiv:2509.02281cs.LGcs.MM2025-09被引 1

提出新方法解决多模态学习中强模态压制弱模态的问题

Balanced Multimodal Learning: An Unidirectional Dynamic Interaction Perspective

  • 采用单向动态交互机制,分步训练主模态并引导其他模态
  • 在多个数据集上超越现有方法,提升多模态任务性能
  • 适合模态差异大、易失衡的多模态场景

多模态学习通常使用联合损失函数融合不同模态以提升模型性能,但这种策略常引发模态失衡——强模态压制弱模态,限制各模态自身信息及跨模态交互信息的充分利用。现有方法如动态损失加权、辅助目标和梯度调制虽能缓解失衡,但均基于联合损失,属于事后修正,未改变联合损失固有的竞争本质。为此,本文提出无监督单向动态交互(UDI)策略,摒弃传统联合损失,采用主动的序列训练方案:先将主模态训练至收敛,再利用其学习到的表示通过无监督损失引导其他模态;动态调整模态间交互,使模型适应任务需求,确保各模态最优贡献。通过解耦模态优化与定向信息流,UDI有效防止单一模态主导,促进跨模态特征学习。实验表明,UDI在处理模态失衡方面显著优于现有方法,提升多模态学习任务性能。

原文摘要 · Abstract (English)

Multimodal learning typically utilizes multimodal joint loss to integrate different modalities and enhance model performance. However, this joint learning strategy can induce modality imbalance, where strong modalities overwhelm weaker ones and limit exploitation of individual information from each modality and the inter-modality interaction information. Existing strategies such as dynamic loss weighting, auxiliary objectives and gradient modulation mitigate modality imbalance based on joint loss. These methods remain fundamentally reactive, detecting and correcting imbalance after it arises, while leaving the competitive nature of the joint loss untouched. This limitation drives us to explore a new strategy for multimodal imbalance learning that does not rely on the joint loss, enabling more effective interactions between modalities and better utilization of information from individual modalities and their interactions. In this paper, we introduce Unidirectional Dynamic Interaction (UDI), a novel strategy that abandons the conventional joint loss in favor of a proactive, sequential training scheme. UDI first trains the anchor modality to convergence, then uses its learned representations to guide the other modality via unsupervised loss. Furthermore, the dynamic adjustment of modality interactions allows the model to adapt to the task at hand, ensuring that each modality contributes optimally. By decoupling modality optimization and enabling directed information flow, UDI prevents domination by any single modality and fosters effective cross-modal feature learning. Our experimental results demonstrate that UDI outperforms existing methods in handling modality imbalance, leading to performance improvement in multimodal learning tasks.

多模态学习模态平衡动态交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。