arXiv:2510.10534cs.CVcs.LG2025-10中稿 · version of an arti…被引 6

解决多模态数据缺失不均衡问题,提升模型鲁棒性。

MCE: Towards a General Framework for Handling Missing Modalities under Imbalanced Missing Rates

  • 动态平衡各模态学习进度,缓解高缺失率模态更新不足。
  • 通过子集预测与跨模态补全,增强特征语义与鲁棒性。
  • 在四个基准上优于现有方法,适合多模态缺失场景应用。

多模态学习在模式识别中取得显著进展,但面对缺失模态,尤其在缺失率不平衡的情况下仍面临挑战。这种不平衡引发恶性循环:缺失率高的模态获得较少更新,导致学习进度不一致和表征退化,进一步削弱其贡献。现有方法多关注全局数据集层面的平衡,忽视样本级模态效用差异及特征质量下降的核心问题。本文提出模态能力增强(MCE)框架,包含两个协同组件:i) 学习能力增强(LCE),引入多层次因素动态平衡模态学习进度;ii) 表示能力增强(RCE),通过子集预测和跨模态补全任务提升特征语义与鲁棒性。在四个多模态基准上的综合评估表明,MCE在多种缺失配置下持续优于当前最优方法。

原文摘要 · Abstract (English)

Multi-modal learning has made significant advances across diverse pattern recognition applications. However, handling missing modalities, especially under imbalanced missing rates, remains a major challenge. This imbalance triggers a vicious cycle: modalities with higher missing rates receive fewer updates, leading to inconsistent learning progress and representational degradation that further diminishes their contribution. Existing methods typically focus on global dataset-level balancing, often overlooking critical sample-level variations in modality utility and the underlying issue of degraded feature quality. We propose Modality Capability Enhancement (MCE) to tackle these limitations. MCE includes two synergistic components: i) Learning Capability Enhancement (LCE), which introduces multi-level factors to dynamically balance modality-specific learning progress, and ii) Representation Capability Enhancement (RCE), which improves feature semantics and robustness through subset prediction and cross-modal completion tasks. Comprehensive evaluations on four multi-modal benchmarks show that MCE consistently outperforms state-of-the-art methods under various missing configurations. The final published version is now available at https://doi.org/10.1016/j.patcog.2025.112591. Our code is available at https://github.com/byzhaoAI/MCE.

多模态缺失数据自适应学习特征增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。