arXiv:2607.22954cs.CLstat.AP2026-07

用大模型自动发现病历中的矛盾信息,提升医疗安全。

Toward Automated Detection of Documentation Inconsistencies in Electronic Health Records

  • 分两阶段用大模型识别病历中的潜在矛盾
  • 发现3460处矛盾,影响近70%住院记录
  • 适合医疗AI、临床质控人员参考

目标:分析通用大语言模型(LLM)在真实出院小结中能识别出哪些内部文档矛盾,并揭示限制其大规模可靠应用的常见失败模式。方法:采用两阶段LLM流程——先用Gemini 2.5 Pro进行开放式候选识别,再用Gemini 2.5 Flash进行上下文验证,处理了3000份随机抽取的MIMIC-IV-Note出院小结。部分输出由临床专家手动审查。结果:该流程共识别出3460个候选矛盾,涉及69.7%的住院记录。代表性矛盾涵盖人口统计、过敏史、操作、诊断、检验、用药及诊疗计划等多个领域,直接影响临床判断或患者安全。专家评审还发现,当验证需时间推理、疾病演变背景或门诊用药惯例时,模型常出现失败。讨论:检测高度依赖上下文,许多标记对需锚定至原始段落和临床领域,判断是否为真矛盾或信息缺失。我们提出一种分级本体论,涵盖严格矛盾与模糊情况,并设计框架按类别、段落、领域和矛盾轴线分类每项标记。结论:本研究建立方法论基础与概念框架,为后续经验证的大规模电子病历矛盾分析提供指引。

原文摘要 · Abstract (English)

Objective: To characterize the kinds of internal documentation inconsistencies a general-domain large language model (LLM) can surface from real-world discharge summaries, and to identify recurring failure modes that limit reliability at scale. Materials and Methods: We applied a two-stage LLM pipeline---open-ended candidate identification (Gemini 2.5 Pro) followed by context-grounded verification (Gemini 2.5 Flash)---to 3,000 randomly sampled MIMIC-IV-Note discharge summaries. A subset of the pipeline output was then reviewed manually by clinical experts. Results: Our pipeline surfaced 3,460 candidate inconsistencies, affecting 69.7% of admissions. Representative examples spanned demographics, allergies, procedures, diagnoses, laboratory, medications, and care-planning domains, with direct implications for clinical reasoning or patient safety. Expert review also revealed recurring failure modes that arise when verification requires temporal reasoning, evolving-diagnosis context, or knowledge of outpatient-prescribing conventions the model does not natively possess. Discussion: Detection is highly context-dependent: many flagged pairs require anchoring each statement to its source section and clinical domain, then assessing whether the conflict reflects a true contradiction or missing context. We propose a graded ontology spanning strict contradiction and ambiguity, with a schema characterizing each flagged case by category, section, domain, and inconsistency axis. Conclusion: This formative study establishes a methodological foundation and conceptual framework to guide subsequent validated, large-scale EHR-inconsistency analysis.

医疗AI病历质量大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。