针对多模态数据采样不均导致的模型偏差,提出动态调整采样量的方法。
Rebalanced Multimodal Learning with Data-aware Unimodal Sampling
- 根据学习状态动态调整每轮采样数量,缓解采样引起的模态不平衡
- 在多个数据集上超越现有最优方法,提升显著
- 可无缝集成到各类多模态模型中,适合作为通用插件使用
为解决由模态不平衡引发的多模态学习退化问题,现有方法主要从模型优化角度平衡各模态训练。然而,几乎所有方法忽略了由单模态数据采样造成的不平衡——等量采样常导致信息量差异,进而引发模态失衡。为此,本文提出一种新型多模态学习方法Data-aware Unimodal Sampling(DUS),旨在动态缓解采样导致的模态不平衡。首先,设计一种累积模态差异度量来监控多模态学习过程;基于学习状态,提出启发式与强化学习(RL)两种数据感知的单模态采样策略,自适应确定每轮采样数量,从而从采样层面缓解不平衡。该方法可无缝嵌入绝大多数现有多模态学习框架作为插件。实验表明,DUS在多个基准上优于多种先进基线,性能达到最佳。
原文摘要 · Abstract (English)
To address the modality learning degeneration caused by modality imbalance, existing multimodal learning~(MML) approaches primarily attempt to balance the optimization process of each modality from the perspective of model learning. However, almost all existing methods ignore the modality imbalance caused by unimodal data sampling, i.e., equal unimodal data sampling often results in discrepancies in informational content, leading to modality imbalance. Therefore, in this paper, we propose a novel MML approach called \underline{D}ata-aware \underline{U}nimodal \underline{S}ampling~(\method), which aims to dynamically alleviate the modality imbalance caused by sampling. Specifically, we first propose a novel cumulative modality discrepancy to monitor the multimodal learning process. Based on the learning status, we propose a heuristic and a reinforcement learning~(RL)-based data-aware unimodal sampling approaches to adaptively determine the quantity of sampled data at each iteration, thus alleviating the modality imbalance from the perspective of sampling. Meanwhile, our method can be seamlessly incorporated into almost all existing multimodal learning approaches as a plugin. Experiments demonstrate that \method~can achieve the best performance by comparing with diverse state-of-the-art~(SOTA) baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。