arXiv:2510.21797cs.LGcs.AI2025-10

通过概率分离动态缓解多模态学习中的样本级不平衡问题。

Mitigating Sample-Level Imbalance via Probabilistic Separation for Adaptive Multimodal Fusion

  • 用模态差距量化单模分支预测差异,识别不平衡样本
  • 基于高斯混合模型实现样本分组,动态调整优化优先级
  • 适合处理多模态数据中存在明显偏差的场景

多模态学习面临模态不平衡问题,主导模态因收敛速度不一致而压制弱模态。现有静态或启发式方法忽略样本级预测偏差差异,无法有效分离低质量异常样本。为此,本文提出一种新型框架,在样本层面定量诊断并动态缓解模态不平衡。首先引入模态差距(Modality Gap)指标,量化单模分支间的预测差异。实证分析揭示其呈现双峰分布,反映平衡与不平衡样本子群的自然共存。随后采用高斯混合模型(GMM)建模该差距分布,利用贝叶斯后验概率实现子群的概率软分离。进一步构建两阶段训练框架:预热阶段与自适应训练阶段。在自适应训练阶段,基于GMM引导的自适应损失动态重分配优化优先级,对不平衡样本施加强调模态对齐惩罚,对平衡样本则优先促进多模态融合。实验表明,本方法显著优于当前最优基线。此外,基于GMM筛选的高质量平衡子集可用于有效数据净化。

原文摘要 · Abstract (English)

Multimodal learning faces modality imbalance, where dominant modalities suppress weaker ones due to inconsistent convergence rates. Existing static or heuristic methods overlook sample-level variations in prediction bias and fail to isolate low-quality outlier samples. To address this, we propose a novel framework to quantitatively diagnose and dynamically mitigate modality imbalance at the sample level. We first introduce a Modality Gap metric to quantify prediction discrepancies between unimodal branches. Empirical analysis reveals a distinct bimodal distribution, reflecting the natural coexistence of balanced and imbalanced sample subgroups. We then employ a Gaussian Mixture Model (GMM) to model this gap distribution, leveraging Bayesian posterior probabilities for probabilistic soft separation of subgroups. Next, we construct a two-stage training framework comprising a Warm-up stage and an Adaptive Training stage. In the Adaptive Training stage, a GMM-guided Adaptive Loss dynamically reallocates optimization priorities, imposing stronger modality alignment penalties on imbalanced samples while prioritizing multimodal fusion for balanced ones. Experimental results demonstrate that our method significantly outperforms current state-of-the-art baselines. Furthermore, fine-tuning on a high-quality balanced subset filtered by the GMM serves as an effective data purification strategy.

多模态学习不平衡自适应融合高斯混合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。