通过改进的模态丢弃与对比学习,提升医疗多模态模型在数据缺失时的诊断能力。
Learning Contrastive Multimodal Fusion with Improved Modality Dropout for Disease Detection and Prediction
- 引入可学习的模态标记,增强对缺失模态的融合感知能力。
- 在仅单模态可用时仍达顶尖性能,验证了强鲁棒性。
- 适配最新CT基础模型,适合临床真实场景部署。
随着医疗诊断日益依赖多模态数据,机器学习模型需有效融合异构信息并具备对缺失模态的鲁棒性。本文提出一种新型多模态学习框架,结合改进的模态丢弃与对比学习,应对实际中模态不平衡与缺失的问题。方法引入可学习的模态标记,提升对缺失模态的融合感知;同时将传统单模态对比目标扩展至融合后的多模态表示。在包含视觉与表格数据的大规模临床数据集上验证,该框架在疾病检测与预测任务中表现卓越,尤其在仅单模态可用的挑战性场景下优势显著。进一步实验表明其可成功集成至近期的CT基础模型。结果证明该方法在有效性、效率与泛化性方面均具优势,是一种可扩展、低成本且极具临床应用潜力的解决方案。代码已公开于https://github.com/omron-sinicx/medical-modality-dropout。
原文摘要 · Abstract (English)
As medical diagnoses increasingly leverage multimodal data, machine learning models are expected to effectively fuse heterogeneous information while remaining robust to missing modalities. In this work, we propose a novel multimodal learning framework that integrates enhanced modalities dropout and contrastive learning to address real-world limitations such as modality imbalance and missingness. Our approach introduces learnable modality tokens for improving missingness-aware fusion of modalities and augments conventional unimodal contrastive objectives with fused multimodal representations. We validate our framework on large-scale clinical datasets for disease detection and prediction tasks, encompassing both visual and tabular modalities. Experimental results demonstrate that our method achieves state-of-the-art performance, particularly in challenging and practical scenarios where only a single modality is available. Furthermore, we show its adaptability through successful integration with a recent CT foundation model. Our findings highlight the effectiveness, efficiency, and generalizability of our approach for multimodal learning, offering a scalable, low-cost solution with significant potential for real-world clinical applications. The code is available at https://github.com/omron-sinicx/medical-modality-dropout.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。