arXiv:2507.03165cs.LG2025-07被引 3

提出PiCME框架,系统评估多模态临床数据的对比学习效果

PiCME: Pipeline for Contrastive Modality Evaluation and Encoding in the MIMIC Dataset

  • 构建多模态对比学习流水线,覆盖五类临床数据组合
  • 三模态表现最优,五模态性能下降但门控LSTM提升显著
  • 兼顾可解释性与公平性,适合临床预测模型优化

多模态深度学习有望通过整合文本、影像、时间序列和结构化人口统计信息等多元患者数据提升临床预测能力。对比学习通过生成可跨任务复用的统一表征,减少对独立模型的需求。尽管在视觉-语言领域取得成功,其在临床中的应用仍主要集中于图像与文本配对。本文提出针对MIMIC数据集的对比模态评估与编码流水线(PiCME),系统评估五类临床数据:出院摘要、放射科报告、胸部X光片、人口统计信息与时间序列数据。我们在所有26种两到五模态组合上预训练对比模型,并在院内死亡率与表型预测任务上评估其效果。为解决模态增多导致的性能瓶颈,引入模态门控LSTM,根据对比学习得到的重要性动态加权各模态。结果表明,对比模型在三模态设置下表现接近监督基线,五模态时性能下降,而监督模型无法恢复;门控LSTM将五模态下的AUROC从73.19%提升至76.93%,AUPRC从51.27%提升至62.26%。我们还对比了对比学习重要性得分与归因分数,并评估了不同人口学子组的泛化能力,凸显其可解释性与公平性优势。PiCME是首个在MIMIC中系统覆盖所有模态组合的对比学习框架,为模态选择、训练策略与公平临床预测提供指导。

原文摘要 · Abstract (English)

Multimodal deep learning holds promise for improving clinical prediction by integrating diverse patient data, including text, imaging, time-series, and structured demographics. Contrastive learning facilitates this integration by producing a unified representation that can be reused across tasks, reducing the need for separate models or encoders. Although contrastive learning has seen success in vision-language domains, its use in clinical settings remains largely limited to image and text pairs. We propose the Pipeline for Contrastive Modality Evaluation and Encoding (PiCME), which systematically assesses five clinical data types from MIMIC: discharge summaries, radiology reports, chest X-rays, demographics, and time-series. We pre-train contrastive models on all 26 combinations of two to five modalities and evaluate their utility on in-hospital mortality and phenotype prediction. To address performance plateaus with more modalities, we introduce a Modality-Gated LSTM that weights each modality according to its contrastively learned importance. Our results show that contrastive models remain competitive with supervised baselines, particularly in three-modality settings. Performance declines beyond three modalities, which supervised models fail to recover. The Modality-Gated LSTM mitigates this drop, improving AUROC from 73.19% to 76.93% and AUPRC from 51.27% to 62.26% in the five-modality setting. We also compare contrastively learned modality importance scores with attribution scores and evaluate generalization across demographic subgroups, highlighting strengths in interpretability and fairness. PiCME is the first to scale contrastive learning across all modality combinations in MIMIC, offering guidance for modality selection, training strategies, and equitable clinical prediction.

多模态学习临床预测对比学习可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。