arXiv:2603.00046cs.LGcs.AI2026-03被引 1

针对医疗多模态数据缺失导致的分布偏斜问题,提出新框架提升模型泛化能力。

REMIND: Rethinking Medical High-Modality Learning under Missingness--A Long-Tailed Distribution Perspective

  • 设计分组专用专家混合架构,为不同模态组合学习专属融合方法。
  • 在真实医疗数据集上显著优于现有方法,尤其对稀有模态组合提升明显。
  • 适合处理高维模态缺失场景,如电子病历、影像与基因数据融合任务。

医疗多模态学习需整合多种异构数据,但在实际临床应用中,因采集限制,患者常缺少完整模态数据,称为‘高模态缺失学习’。我们发现,这种缺失导致模态组合数量呈指数增长,并形成长尾分布,而以往研究忽视了这一现象,致使尾部组合性能严重下降。分析表明,根本原因在于:1)梯度不一致,尾部组合更新方向偏离全局优化;2)概念漂移,每种模态组合需独立融合策略。为此,提出REMIND框架,从长尾分布视角重思高模态缺失下的多模态学习。核心是采用新型分组专用专家混合架构,可扩展地学习任意模态组合的融合函数,同时结合分布鲁棒优化策略,对低频组合进行加权。在多个真实医疗数据集上的实验表明,该框架持续超越现有最优方法,在高模态缺失下具有强泛化能力。

原文摘要 · Abstract (English)

Medical multi-modal learning is critical for integrating information from a large set of diverse modalities. However, when leveraging a high number of modalities in real clinical applications, it is often impractical to obtain full-modality observations for every patient due to data collection constraints, a problem we refer to as 'High-Modality Learning under Missingness'. In this study, we identify that such missingness inherently induces an exponential growth in possible modality combinations, followed by long-tail distributions of modality combinations due to varying modality availability. While prior work overlooked this critical phenomenon, we find this long-tailed distribution leads to significant underperformance on tail modality combination groups. Our empirical analysis attributes this problem to two fundamental issues: 1) gradient inconsistency, where tail groups' gradient updates diverge from the overall optimization direction; 2) concept shifts, where each modality combination requires distinct fusion functions. To address these challenges, we propose REMIND, a unified framework that REthinks MultImodal learNing under high-moDality missingness from a long-tail perspective. Our core idea is to propose a novel group-specialized Mixture-of-Experts architecture that scalably learns group-specific multi-modal fusion functions for arbitrary modality combinations, while simultaneously leveraging a group distributionally robust optimization strategy to upweight underrepresented modality combinations. Extensive experiments on real-world medical datasets show that our framework consistently outperforms state-of-the-art methods, and robustly generalizes across various medical multi-modal learning applications under high-modality missingness.

多模态学习医疗AI数据缺失长尾分布

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。