arXiv:2505.16941cs.LGcs.AI2025-05被引 9

首个针对电子病历基础模型的临床意义评估基准,验证其在真实医疗场景中的表现。

FoMoH: A clinically meaningful foundation model evaluation for structured electronic health records

  • 构建14项临床预测任务,涵盖急慢性病早期诊断与预后
  • 发现顶尖模型在少标注数据下仍具强区分能力,且跨人群更公平
  • 揭示模型在低发病率疾病上表现差、跨机构迁移难,适合临床研究者参考

基础模型(FMs)有望解决传统监督学习的三大痛点:对大量标注数据的依赖、任务特异性及泛化能力差。尽管结构化电子健康记录(EHR)基础模型取得进展,尚无系统性基准验证其是否真正实现这些承诺。本文提出涵盖患者预后与急慢性病早期诊断的14项临床有意义预测任务,基于超过600万例来自哥伦比亚大学医学院和MIMIC-IV的数据,对6个先进EHR FMs进行评估,严格控制数据污染并确保可复现性。结果表明:顶级模型在有限标注数据下优于传统基线,在区分度和跨群体公平性方面表现更优;但低发病率疾病的判别性能下降,标注数据少时校准性较差,跨机构迁移仍存挑战。研究为理解EHR FMs临床价值提供依据,揭示关键差距,并建立可复现的评估框架。

原文摘要 · Abstract (English)

Foundation models (FMs) promise to address core limitations of traditional supervised machine learning: (i) reliance on large amounts of labeled data, (ii) task specificity, and (iii) poor transportability. Despite methodological advances in structured electronic health record (EHR) foundation models, no systematic benchmark has validated whether these models meaningfully deliver on these promises. We introduce a benchmark of 14 clinically meaningful prediction tasks spanning patient prognosis and early diagnosis of acute and chronic conditions. We benchmark 6 state-of-the-art EHR FMs beyond population-level discrimination, emphasizing the need for evaluating their calibration and fairness, with rigorous controls for data contamination and reproducibility across more than 6 million patients from Columbia University Irving Medical Center and MIMIC-IV. Our benchmark identifies that FMs deliver on some of their promises. In particular, top-performing FMs outperform traditional baselines on discriminative performance, especially under limited labeled data, and exhibit more equitable performance across socio-medical groups. However, these models may underperform in low-prevalence settings, as pretraining losses may discard discriminative information about such conditions, and present lower calibration under limited labeled data. Further, cross-institutional transportability remains a challenge for structured EHR FMs. Together, these findings advance our understanding of EHR FMs' potential for clinical utility, highlight critical gaps that remain to be addressed, and provide a reproducible framework to track progress.

电子病历基础模型临床评估公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。