MedMIX融合多专家模型,提升医学多模态诊断在缺失数据下的稳定性。
MedMIX: Modality-Internal Expert Fusion for Multimodal Medical Diagnosis

- 每模态内聚合多个小模型的互补特征,跨模态学习动态融合策略。
- 在三个基准上均表现优异,缺失模态时性能下降不超过5%。
- 适合医疗场景中数据不全、模态差异大的真实应用。
多模态临床预测面临三大挑战:每模态需多个基础模型(FMs)发挥互补优势,训练与测试时普遍存在模态缺失,且不同样本对模态依赖程度各异。我们提出MedMIX框架,结合模态内专家融合、可学习的跨模态融合以及仅训练阶段的大-小模型协作,以增强在不完备模态下的鲁棒性。每个模态内,MedMIX融合多个小型专家模型的嵌入表示;跨模态时,对可用模态进行自适应融合;训练阶段利用大教师模型优化部署表示,无额外推理开销。在三个异构基准(OpenI、MIMIC-IV-MM 和 MMIST-ccRCC)上,MedMIX表现持续领先,并在受控模态缺失扰动下保持稳定;进一步在MIMIC-III上验证了跨队列迁移下的持续鲁棒性。结果表明,MedMIX是一个统一整合模态内专家协作、样本级跨模态融合与高效大-小模型协同的实用框架,对模态缺失具有强鲁棒性。
原文摘要 · Abstract (English)
Multimodal clinical prediction faces three challenges: multiple foundation models (FMs) with complementary strengths per modality, pervasive missing modalities at training and test time, and sample-specific variation in modality contributions. We introduce MedMIX, a multimodal framework that combines intra-modality expert fusion, learned inter-modality fusion, and training-only large--small model collaboration for robust medical prediction under incomplete modalities. Within each modality, MedMIX aggregates complementary embeddings from multiple small expert models; across modalities, it performs learned fusion over available modalities; and during training, it leverages large teacher models to improve deployed representations without additional inference cost. Across three heterogeneous benchmarks (OpenI, MIMIC-IV-MM, and MMIST-ccRCC), MedMIX achieves consistently strong performance while remaining robust under controlled missing-modality perturbations, and further demonstrates sustained robustness under cross-cohort shift on MIMIC-III. These results highlight MedMIX as a practical framework that unifies within-modality expert collaboration, sample-specific cross-modality fusion, and efficient large--small model collaboration while remaining robust to incomplete modalities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。