arXiv:2506.04156cs.CL2025-06被引 19

构建首个面向患者住院信息需求的临床病历问答数据集,助力AI生成准确可信的回答。

A Dataset for Addressing Patient's Information Needs related to Clinical Course of Hospitalization

  • 基于真实重症与急诊病例,标注患者提问、临床笔记及专家答案
  • 134个病例测试显示,先给答案再引用文献的提示策略表现最佳
  • 适用于医疗AI研发者和临床信息系统的评估,推动可信赖的患者问答系统发展

患者在住院期间有特定的信息需求,可通过电子健康记录(EHR)中的临床证据满足。尽管人工智能(AI)系统在回应这些需求方面展现潜力,但评估其回答的真实性与相关性仍需高质量数据集。目前尚无公开数据集能捕捉患者在真实EHR背景下的信息需求。我们提出ArchEHR-QA,一个由专家标注的真实患者案例数据集,涵盖重症监护室与急诊科场景。每个案例包含患者在公共健康论坛提出的疑问、医生解读后的对应问题、相关的临床笔记片段(含句级相关性标注)以及医生撰写的答案。为建立基于EHR的问答基准,我们评估了三种开源大模型(Llama 4、Llama 3、Mixtral)在三种提示策略下的表现:(1)生成带临床句子引用的答案,(2)先生成答案再引用,(3)从筛选后的引用中生成答案。评估维度包括事实性(引用句子与真实答案的重合度)和相关性(系统回答与参考答案在文本与语义上的相似度)。最终数据集包含134个患者案例。结果表明,答案先行的提示策略表现最优,其中Llama 4得分最高。人工错误分析验证了该结果,并揭示常见问题如遗漏关键临床证据、矛盾或虚构内容。总体而言,ArchEHR-QA为开发与评估以患者为中心的EHR问答系统提供了有力基准,凸显了在临床场景中生成事实准确且相关回答的迫切需求。

原文摘要 · Abstract (English)

Patients have distinct information needs about their hospitalization that can be addressed using clinical evidence from electronic health records (EHRs). While artificial intelligence (AI) systems show promise in meeting these needs, robust datasets are needed to evaluate the factual accuracy and relevance of AI-generated responses. To our knowledge, no existing dataset captures patient information needs in the context of their EHRs. We introduce ArchEHR-QA, an expert-annotated dataset based on real-world patient cases from intensive care unit and emergency department settings. The cases comprise questions posed by patients to public health forums, clinician-interpreted counterparts, relevant clinical note excerpts with sentence-level relevance annotations, and clinician-authored answers. To establish benchmarks for grounded EHR question answering (QA), we evaluated three open-weight large language models (LLMs)--Llama 4, Llama 3, and Mixtral--across three prompting strategies: generating (1) answers with citations to clinical note sentences, (2) answers before citations, and (3) answers from filtered citations. We assessed performance on two dimensions: Factuality (overlap between cited note sentences and ground truth) and Relevance (textual and semantic similarity between system and reference answers). The final dataset contains 134 patient cases. The answer-first prompting approach consistently performed best, with Llama 4 achieving the highest scores. Manual error analysis supported these findings and revealed common issues such as omitted key clinical evidence and contradictory or hallucinated content. Overall, ArchEHR-QA provides a strong benchmark for developing and evaluating patient-centered EHR QA systems, underscoring the need for further progress toward generating factual and relevant responses in clinical contexts.

医疗AI问答系统EHR数据大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。