通过置信度引导的多阶段融合,提升医疗多模态数据预测性能。
MedPatch: Confidence-Guided Multi-Stage Fusion for Multimodal Clinical Data
- 分阶段融合:联合与延迟融合并用,动态整合多源医疗数据。
- 在MIMIC-IV等数据集上,死亡率预测准确率超越现有模型。
- 适合处理缺失模态的临床数据,对医疗AI研究者有参考价值。
临床决策依赖于多种数据模态的融合,如临床时间序列、医学影像和文本报告。与其它领域相比,真实医疗数据具有异构性、数据量有限且因模态缺失而稀疏,严重制约模型在临床预测任务中的表现。受临床工作流程启发,我们提出MedPatch,一种多阶段多模态融合架构,通过置信度引导的分块整合实现多模态无缝融合。MedPatch包含三个核心组件:(i) 多阶段融合策略,同时利用联合融合与延迟融合;(ii) 缺失感知模块,处理缺失模态的稀疏样本;(iii) 联合融合模块,基于校准后的单模态令牌置信度对潜在令牌块进行聚类。我们在MIMIC-IV、MIMIC-CXR和MIMIC-Notes数据集上,使用临床时间序列、胸部X光片、放射科报告和出院记录,评估了MedPatch在住院死亡率预测和临床状态分类两个基准任务上的表现。相比现有基线模型,MedPatch达到当前最优性能。本工作验证了置信度引导的多阶段融合在应对多模态数据异构性方面的有效性,并为临床预测任务建立了新的基准结果。
原文摘要 · Abstract (English)
Clinical decision-making relies on the integration of information across various data modalities, such as clinical time-series, medical images and textual reports. Compared to other domains, real-world medical data is heterogeneous in nature, limited in size, and sparse due to missing modalities. This significantly limits model performance in clinical prediction tasks. Inspired by clinical workflows, we introduce MedPatch, a multi-stage multimodal fusion architecture, which seamlessly integrates multiple modalities via confidence-guided patching. MedPatch comprises three main components: (i) a multi-stage fusion strategy that leverages joint and late fusion simultaneously, (ii) a missingness-aware module that handles sparse samples with missing modalities, (iii) a joint fusion module that clusters latent token patches based on calibrated unimodal token-level confidence. We evaluated MedPatch using real-world data consisting of clinical time-series data, chest X-ray images, radiology reports, and discharge notes extracted from the MIMIC-IV, MIMIC-CXR, and MIMIC-Notes datasets on two benchmark tasks, namely in-hospital mortality prediction and clinical condition classification. Compared to existing baselines, MedPatch achieves state-of-the-art performance. Our work highlights the effectiveness of confidence-guided multi-stage fusion in addressing the heterogeneity of multimodal data, and establishes new state-of-the-art benchmark results for clinical prediction tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。