构建急诊临床推理真实数据集,评估大模型随证据累积的判断能力
ER-Reason: A Benchmark Dataset for LLM Clinical Reasoning in the Emergency Room
- 基于2.5万条真实急诊病历,覆盖从分诊到诊断全流程任务
- 用2555位医生标注的渐进式判断题,衡量模型信念更新方向与幅度
- 突破传统题库局限,揭示大模型在真实场景下的推理缺陷
现有大语言模型临床推理评估基准普遍存在定义模糊、任务单一、依赖虚构病例等问题,导致其在真实临床研究中表现与考试成绩差异显著。为此,我们提出ER-Reason,一个面向急诊医学全流程的基准数据集,包含3,437名患者共25,174条去标识化临床记录,涵盖分诊、治疗选择、转归规划和最终诊断四个阶段。评估不仅关注诊断准确率,更引入基于真实病例的逐步式脚本一致性测试(SCT)问题,通过2,555名急诊医师标注,衡量模型在证据积累过程中是否以正确方向与程度更新诊断信念。我们在该数据集上评估了推理型与非推理型大模型,结果表明,该任务能更细致揭示模型在真实病例中的推理短板。
原文摘要 · Abstract (English)
Existing benchmarks for evaluating the clinical reasoning capabilities of large language models (LLMs) often lack a clear definition of "clinical reasoning" as a construct, fail to capture the full breadth of interdependent tasks within a clinical workflow, and rely on stylized vignettes rather than real-world clinical documentation. As a result, recent studies have found significant discrepancies between LLM performance on stylized benchmarks derived from medical licensing exams and their performance in real-world prospective studies. To address these limitations, we introduce ER-Reason, a benchmark designed to evaluate LLM reasoning as clinical evidence accumulates across decision-making tasks spanning the full workflow of emergency medicine. ER-Reason comprises 25,174 de-identified clinical notes from 3,437 patients, supporting evaluation across all stages of the emergency department workflow: triage intake, treatment selection, disposition planning, and final diagnosis. Crucially, evaluation in ER-Reason extends beyond diagnostic accuracy to include stepwise Script Concordance Test (SCT)-style questions grounded in real patient cases, which assess whether LLMs update their diagnostic beliefs in the correct direction and magnitude as clinical evidence accumulates, scored against 2,555 emergency physician annotations. We evaluate reasoning and non-reasoning LLMs on ER-Reason, and show that our tasks provide a more nuanced view of how LLM reasoning fails on real patient cases than existing benchmarks allow.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。