arXiv:2507.14824cs.LGcs.AI2025-07

用MIMIC-IV数据集评估多模态医疗AI模型表现,验证融合多源数据可提升预测效果。

Benchmarking Foundation Models with Multimodal Public Electronic Health Records

  • 基于MIMIC-IV构建标准化数据流程,统一异构临床记录
  • 8个基础模型对比显示多模态输入显著提升预测性能
  • 结果支持可信多模态AI在真实临床场景的应用

基础模型已成为处理电子健康记录(EHR)的有力方法,能灵活应对多样化的医疗数据模态。本研究提出一个全面基准,利用公开可用的MIMIC-IV数据库,评估基础模型在预测性能、公平性和可解释性方面的表现,涵盖单模态编码器与多模态学习者,包括领域专用与通用变体。为确保评估的一致性与可复现性,我们开发了标准化数据处理管道,将异构临床记录转化为分析就绪格式。系统比较了八种基础模型,结果表明融合多种数据模态可稳定提升预测性能,且未引入额外偏差。该基准旨在推动可信赖多模态人工智能系统在真实临床场景中的发展。代码已开源:https://github.com/nliulab/MIMIC-Multimodal。

原文摘要 · Abstract (English)

Foundation models have emerged as a powerful approach for processing electronic health records (EHRs), offering flexibility to handle diverse medical data modalities. In this study, we present a comprehensive benchmark that evaluates the performance, fairness, and interpretability of foundation models, both as unimodal encoders and as multimodal learners, using the publicly available MIMIC-IV database. To support consistent and reproducible evaluation, we developed a standardized data processing pipeline that harmonizes heterogeneous clinical records into an analysis-ready format. We systematically compared eight foundation models, encompassing both unimodal and multimodal models, as well as domain-specific and general-purpose variants. Our findings demonstrate that incorporating multiple data modalities leads to consistent improvements in predictive performance without introducing additional bias. Through this benchmark, we aim to support the development of effective and trustworthy multimodal artificial intelligence (AI) systems for real-world clinical applications. Our code is available at https://github.com/nliulab/MIMIC-Multimodal.

多模态医疗AIMIMIC-IV基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。