构建百万级真实病历问答数据集,评估大模型临床决策能力
EHRBench: An Automated and Reliable EHR-based Benchmark for Clinical Decision Making with LLMs

- 通过病历-大模型-知识库联动自动构建高质量临床问答数据
- 覆盖诊断、治疗、预后三类任务,共96万条标准化问答
- 为医疗大模型可靠性评估提供可复现的基准测试平台
临床决策是真实医疗流程的核心,医生需在证据不全情况下推断诊断、选择治疗或预测健康结局。尽管大语言模型因强大的语言能力、广泛的生物医学知识和高效率被越来越多用于辅助决策,但其在真实临床任务中的可靠性仍不明确。理想的临床决策评估基准应具备自动化与高可靠性,确保规模与质量并重。为此,我们提出EHRBench,一个基于真实电子病历(EHR)的自动化、可靠的大规模临床决策评估基准。通过构建EHR-LLM-KB(知识库)交互管道,使用专用大模型将就诊记录转化为结构化模板,并确定性生成问答对;同时结合系统化的知识库验证与增强,过滤幻觉或模糊关系,提升数据可靠性。该方法生成近100万(960,067)条涵盖诊断、治疗、预后三类需推理的临床决策任务的问答数据。我们在EHRBench上对30多个代表性大模型进行评测,分析其性能与鲁棒性,结果显示各模型表现趋势一致,进一步验证了基准的可靠性,并揭示了向临床可用大模型系统迈进的关键差距。
原文摘要 · Abstract (English)
Clinical decision-making (CDM) is central to real-world clinical workflows, where clinicians infer diagnoses, select treatments, or anticipate future health outcomes under incomplete evidence. LLMs are increasingly used to support these decisions due to strong language capabilities, broad biomedical knowledge, and efficiency, yet the reliability of LLMs on real-world clinical decision tasks remains insufficiently understood. To evaluate CDM models, especially LLM-based models, an ideal and practical medical decision benchmark should be constructed via an automated yet reliable pipeline to ensure both scale and quality. Moreover, the grounding of a CDM benchmark in real patient EHRs can better support evaluation on practical CDM tasks that require substantive biomedical knowledge and clinical inference. To fill the gaps, we introduce EHRBench, an automated and reliable EHR-grounded benchmark for evaluating LLM-based clinical decision-making at scale. To ensure scalability and reliability, EHRBench is constructed through an EHR-LLM-KB(knowledge-base) interaction pipeline. For efficiency, we use a specialized LLM to automatically convert encounter-level EHR trajectories into structured templates and deterministically instantiate the templates into QA items. In parallel, we apply systematic KB-based verification and enrichment to filter hallucinated or ambiguous relations and to improve reliability. Using this pipeline, we construct nearly 1M (960,067) QA items spanning three core inference-required clinical decision tasks: diagnosis, treatment, and prognosis. We benchmark more than 30 representative LLMs on EHRBench and provide detailed analyses of performance and robustness. The results show consistent capability trends across settings, further validating the reliability of EHRBench and highlighting actionable gaps toward clinically reliable LLM systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。