arXiv:2605.30295cs.CLcs.AI2026-05中稿 · ICML被引 1

构建首个面向临床真实场景的结构化病历评测数据集

MedCase-Structured: A Text-to-FHIR Dataset for Benchmarking Diagnostic Reasoning in Clinically Realistic EHR Settings

论文配图:MedCase-Structured: A Text-to-FHIR Dataset for Benchmarking Diagnostic Reasoning in Clinically Realistic EHR Settings
图 1 · 摘自论文原文
  • 用大模型从文本生成符合医疗标准的结构化病历
  • 97.1%生成的病历符合标准且无虚构编码
  • 适合评估大模型在真实医疗系统中的诊断能力

大型语言模型在临床推理与决策支持方面展现出潜力,但其在结构化电子病历(EHR)环境下的评估仍有限。现有基准多依赖静态数据集或非结构化输入,无法反映临床系统中使用的互操作性数据格式。本文提出一种可复用的生成管道,将非结构化文本转化为符合HL7 FHIR R4标准的病历包,支持对临床决策支持系统的可控评估。该管道结合分阶段大模型生成与术语校验修复机制,消除虚构编码并保证结构与语义一致性。基于MedCaseReasoning数据集,构建了包含1,732个FHIR包的MedCase-Structured合成数据集,97.1%的病例成功生成完整有效包。在该数据集上的评估显示,大模型在结构化FHIR输入下的诊断准确率显著低于纯文本输入,凸显部署对齐评估的重要性。

原文摘要 · Abstract (English)

Large language models (LLMs) show promise for clinical reasoning and decision support, but evaluation in structured, electronic health record-congruent settings remains limited. Existing benchmarks often rely on static datasets or unstructured inputs that do not reflect the interoperable data formats used in clinical systems. We introduce a reusable pipeline for generating terminology-grounded HL7 FHIR R4 bundles from unstructured text, enabling controllable evaluation of clinical decision support systems over structured inputs. The pipeline combines staged LLM generation with terminology-grounded validation and repair to eliminate hallucinated codes and enforce structural and semantic consistency. Applying this approach to MedCaseReasoning, we construct MedCase-Structured, a synthetic dataset of 1,732 FHIR bundles derived from clinician-authored diagnostic cases, producing complete, valid bundles for 97.1% of attempted cases. Evaluation on MedCase-Structured reveals consistently lower diagnostic accuracy for LLMs on structured FHIR inputs than with plain text, highlighting the importance of deployment-aligned benchmarking.

医疗AIFHIR大模型评测结构化数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。