让病人群体成为AI模型的核心,提升医疗AI的可解释性与可信度。
Cohort-Anchored Foundation Models for Electronic Health Records: From Risk Scores to Auditable Peer Cohorts

- 以病人群体为中心设计模型,从数据到训练全程锚定临床群体结构
- 在四项真实临床任务中显著提升风险预测与报告生成效果
- 适合关注可审计、可解释医疗AI的研究者与临床开发者
基础模型在医学问答、影像和电子健康记录(EHR)任务中表现卓越,但临床部署仍受限于可解释性差、对分布偏移敏感以及与医生推理不一致等问题。我们指出,现有方法过于注重表征学习,将患者比较视为衍生属性而非核心证据。为此,提出Cohort-Anchored Foundation Model(CAFM)框架,将病人群体作为学习流程中的首要对象。该框架包含四个阶段:偏差感知的数据清洗、基于病人群体的预训练、多模态病人群体对齐、医生参与的迭代优化。该方法提升数据质量,围绕临床有意义的群体结构组织表征,保留模态特异性关系,并支持可审计的临床决策。框架具备组合性,无需修改原有EHR基础模型编码器即可增强。通过四项临床案例验证:急性肾损伤预测、心电图心血管风险分层、眼眶影像视神经病变分诊、电生理报告生成。同时提出五个可验证假设,并识别数据质量、时间不规则性、多模态学习、分布偏移及超越预测准确率的评估等开放挑战。我们认为,明确以病人群体锚定基础模型,是迈向可信临床AI的系统性路径。
原文摘要 · Abstract (English)
Foundation models have achieved remarkable performance across medical question answering, imaging, and electronic health record (EHR) tasks, yet reliable clinical deployment remains challenging due to limited interpretability, vulnerability to distribution shift, and weak alignment with clinician reasoning. We argue that these limitations arise because existing approaches prioritize representation learning while treating patient comparison as an emergent property rather than a primary source of clinical evidence. To address this gap, we propose CAFM, a Cohort-Anchored Foundation Model framework that elevates patient cohorts to a first-class object throughout the learning pipeline. The framework consists of four stages: deviation-aware data curation, cohort-conditioned pretraining, multimodal cohort alignment, and clinician-in-the-loop refinement. Together, these stages improve data quality, organize representations around clinically meaningful cohort structure, preserve modality-specific relationships, and support auditable clinical decision-making. The framework is compositional and can augment existing EHR foundation models without modifying their underlying encoders. We illustrate CAFM through four clinical case studies spanning acute kidney injury prediction, cardiovascular risk stratification from electrocardiograms, optic neuropathy triage from orbital imaging, and electroretinogram-grounded report generation. We further present five empirically testable hypotheses and identify open challenges in data quality, irregular temporality, multimodal learning, distribution shift, and evaluation beyond predictive accuracy. We argue that explicitly anchoring foundation models to patient cohorts provides a principled path toward trustworthy clinical AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。