arXiv:2506.11550cs.LGcs.AI2025-06ICML被引 10

通过数据重混解决多模态学习不平衡与不足问题

Improving Multimodal Learning Balance and Sufficiency through Data Remixing

  • 分离多模态数据并筛选难样本,缓解模态失衡
  • 批量重装对齐梯度方向,提升单模态充分性
  • 无需额外计算开销,显著提升模型准确率

不同模态在优化轨迹上存在显著差异,包括速度和路径,导致联合训练时出现模态懒惰与模态冲突,造成多模态学习不充分且不平衡。现有方法主要通过强化弱模态、对齐优化速度或分解学习来改善,但难以同时实现单模态充分性和多模态平衡。本文首次提出多模态数据重混(Data Remixing):先解耦多模态数据并为各模态过滤难样本以缓解模态不平衡;再进行批次级重装,对齐梯度方向,避免跨模态干扰,从而增强单模态学习充分性。实验表明,该方法可无缝集成至现有框架,在CREMAD上提升准确率约6.50%,在Kinetic-Sounds上提升3.41%,无需训练集扩展或推理阶段额外计算开销。

原文摘要 · Abstract (English)

Different modalities hold considerable gaps in optimization trajectories, including speeds and paths, which lead to modality laziness and modality clash when jointly training multimodal models, resulting in insufficient and imbalanced multimodal learning. Existing methods focus on enforcing the weak modality by adding modality-specific optimization objectives, aligning their optimization speeds, or decomposing multimodal learning to enhance unimodal learning. These methods fail to achieve both unimodal sufficiency and multimodal balance. In this paper, we, for the first time, address both concerns by proposing multimodal Data Remixing, including decoupling multimodal data and filtering hard samples for each modality to mitigate modality imbalance; and then batch-level reassembling to align the gradient directions and avoid cross-modal interference, thus enhancing unimodal learning sufficiency. Experimental results demonstrate that our method can be seamlessly integrated with existing approaches, improving accuracy by approximately 6.50%$\uparrow$ on CREMAD and 3.41%$\uparrow$ on Kinetic-Sounds, without training set expansion or additional computational overhead during inference. The source code is available at https://github.com/MatthewMaxy/Remix_ICML2025.

多模态学习数据重混优化平衡模型性能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。