arXiv:2509.25831cs.LG2025-09NeurIPS被引 3

通过制造语义不一致的跨模态数据,强迫模型学习弱模态信息。

MIDAS: Misalignment-based Data Augmentation Strategy for Imbalanced Multimodal Learning

  • 生成跨模态语义矛盾的数据,用单模态置信度标注。
  • 动态提升低置信度模态损失权重,缓解模态依赖。
  • 优先训练难分辨的混淆样本,适合多模态不平衡场景。

多模态模型常过度依赖主导模态,难以达到最优性能。现有方法多聚焦于优化目标或训练过程,而数据层面的解决方案仍待探索。本文提出MIDAS,一种基于错位的数据增强策略:生成包含语义不一致跨模态信息的样本,并利用单模态置信度进行标注,以迫使模型学习矛盾信号。然而,这种置信度标注仍可能偏向更自信的模态。为此,我们引入弱模态加权机制,动态提高置信度最低模态的损失权重,从而帮助模型充分使用弱模态。此外,当错位特征与对齐特征相似度更高时,这些样本更具挑战性,有助于模型更好区分类别。为此,我们提出困难样本加权策略,优先处理此类语义模糊的错位样本。在多个多模态分类基准上实验表明,MIDAS显著优于相关基线方法,在缓解模态不平衡方面表现优异。

原文摘要 · Abstract (English)

Multimodal models often over-rely on dominant modalities, failing to achieve optimal performance. While prior work focuses on modifying training objectives or optimization procedures, data-centric solutions remain underexplored. We propose MIDAS, a novel data augmentation strategy that generates misaligned samples with semantically inconsistent cross-modal information, labeled using unimodal confidence scores to compel learning from contradictory signals. However, this confidence-based labeling can still favor the more confident modality. To address this within our misaligned samples, we introduce weak-modality weighting, which dynamically increases the loss weight of the least confident modality, thereby helping the model fully utilize weaker modality. Furthermore, when misaligned features exhibit greater similarity to the aligned features, these misaligned samples pose a greater challenge, thereby enabling the model to better distinguish between classes. To leverage this, we propose hard-sample weighting, which prioritizes such semantically ambiguous misaligned samples. Experiments on multiple multimodal classification benchmarks demonstrate that MIDAS significantly outperforms related baselines in addressing modality imbalance.

多模态学习数据增强模态平衡

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。