首个面向多轮病历摘要的医学问答基准,评估AI是否真能基于证据推理。
EHRNote-ChatQA: A Benchmark for Evidence-Grounded Multi-Turn Clinical Question Answering over Longitudinal Discharge Summaries

- 构建多轮问答模板,结合专家标注与LLM生成,确保答案有据可循。
- 包含16,072对问答,覆盖8类临床场景,每例最多5份出院记录。
- 揭示大模型在多轮推理中易失真,证据追溯能力远弱于单轮问答。
出院摘要包含患者住院全过程信息,是医疗决策的关键依据。医生常需跨多份摘要迭代整合信息并验证证据。尽管大语言模型(LLMs)被用于临床问答,现有基准多聚焦单轮或考试式知识题,缺乏对证据支撑的严格评估。本文提出EHRNote-ChatQA,首个基于纵向出院摘要的证据驱动多轮临床问答基准。数据源自去标识化的MIMIC-IV,包含967个患者级样本(每例1–5份摘要),共16,072对专家验证的问答对(8,036个内容问题,每个配一个证据溯源问题),覆盖8类临床领域。通过专家指导的流程:结构化摘要模板、定制多轮问答框架、LLM生成后经11位医学专家逐项审查修订。对22个开源与闭源模型的评测显示:模型在证据追溯上表现显著差于内容回答;多轮错误会累积;单轮表现无法可靠迁移至多轮场景。该基准为临床问答系统提供严谨评估工具。数据将通过PhysioNet凭证访问公开。
原文摘要 · Abstract (English)
Discharge summaries are crucial clinical documents containing the context of a patient's overall hospital stay, and are routinely reviewed by medical experts for patient readmission, ongoing care, and diagnostic decision-making. When reviewing them, medical experts often must iteratively synthesize information across multiple summaries while verifying the evidence supporting each answer. Although large language models (LLMs) are increasingly explored for clinical question answering, existing benchmarks do not sufficiently reflect this setting: they often evaluate exam-style medical knowledge or focus on single-turn question answering with limited evidence-grounding evaluation. We introduce EHRNote-ChatQA, the first benchmark for evidence-grounded multi-turn clinical question answering over patients' multiple discharge summaries. Built from de-identified MIMIC-IV discharge summaries, EHRNote-ChatQA contains 967 patient-level multi-turn samples spanning one to five notes and 16,072 medical-expert-verified QA pairs (8,036 content questions, each paired with an evidence-grounding question) across eight clinical categories. The benchmark is constructed through an expert-informed pipeline combining discharge-summary structuring schema, expert-curated multi-turn QA templates, and LLM-based generation, followed by review and revision of every single QA sample by 11 medical experts. Benchmarking 22 open- and closed-source LLMs reveals several challenges, including that LLMs struggle more with evidence grounding than content answering, multi-turn errors compound across turns, and single-turn clinical QA performance does not reliably transfer to this setting. These findings establish EHRNote-ChatQA as a rigorous and practical benchmark for evaluating clinical QA systems. The dataset will be made publicly available through PhysioNet credentialed access.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。