构建首个面向临床真实场景的结构化病历评测数据集
MedCase-Structured: A Text-to-FHIR Dataset for Benchmarking Diagnostic Reasoning in Clinically Realistic EHR Settings

- 用大模型从文本生成符合医疗标准的结构化病历
- 97.1%生成的病历符合标准且无虚构编码
- 适合评估大模型在真实医疗系统中的诊断能力
大型语言模型在临床推理与决策支持方面展现出潜力,但其在结构化电子病历(EHR)环境下的评估仍有限。现有基准多依赖静态数据集或非结构化输入,无法反映临床系统中使用的互操作性数据格式。本文提出一种可复用的生成管道,将非结构化文本转化为符合HL7 FHIR R4标准的病历包,支持对临床决策支持系统的可控评估。该管道结合分阶段大模型生成与术语校验修复机制,消除虚构编码并保证结构与语义一致性。基于MedCaseReasoning数据集,构建了包含1,732个FHIR包的MedCase-Structured合成数据集,97.1%的病例成功生成完整有效包。在该数据集上的评估显示,大模型在结构化FHIR输入下的诊断准确率显著低于纯文本输入,凸显部署对齐评估的重要性。
原文摘要 · Abstract (English)
Large language models (LLMs) show promise for clinical reasoning and decision support, but evaluation in structured, electronic health record-congruent settings remains limited. Existing benchmarks often rely on static datasets or unstructured inputs that do not reflect the interoperable data formats used in clinical systems. We introduce a reusable pipeline for generating terminology-grounded HL7 FHIR R4 bundles from unstructured text, enabling controllable evaluation of clinical decision support systems over structured inputs. The pipeline combines staged LLM generation with terminology-grounded validation and repair to eliminate hallucinated codes and enforce structural and semantic consistency. Applying this approach to MedCaseReasoning, we construct MedCase-Structured, a synthetic dataset of 1,732 FHIR bundles derived from clinician-authored diagnostic cases, producing complete, valid bundles for 97.1% of attempted cases. Evaluation on MedCase-Structured reveals consistently lower diagnostic accuracy for LLMs on structured FHIR inputs than with plain text, highlighting the importance of deployment-aligned benchmarking.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。