用大模型自动检查出院记录,提升医疗转诊效率
CareTransition-Audit: A Benchmark to Audit Discharge Summaries for Efficient Care Transitions

- 将临床框架转为46个问题清单,用LLM自动评估出院记录
- 最佳模型完成度达74.2%,与医生判断中等一致(kappa≈0.5)
- 发现模型难识别模糊信息,适合医疗质量改进研究者使用
出院记录不完整或不一致导致医疗碎片化和可避免的再入院。尽管其对患者安全至关重要,当前审计仍依赖人工且无法扩展。本文提出一种基于大语言模型(LLMs)的自动化出院记录审计框架。将DISCHARGED框架转化为包含46个问题的检查清单,基于MIMIC-IV数据库中的50份出院记录(含临床医生标注真值),评估11种LLM。模型评估的平均文档完整性在54.9%至74.2%之间,表现最佳的模型与临床医生标注的Cohen's kappa值约为0.5,表明中等一致性。所有模型均难以识别模糊记录(Unclear),凸显当前自动化审计的关键短板。本工作提供经临床验证的基准与零样本基线,推动临床文档质量系统性提升。
原文摘要 · Abstract (English)
Incomplete or inconsistent discharge documentation drives care fragmentation and avoidable readmissions. Despite its critical role in patient safety, auditing discharge summaries relies on manual review and does not scale. We propose an automated framework for auditing discharge summaries using large language models (LLMs). Our approach operationalizes the DISCHARGED framework into a checklist of 46 questions. Using 50 summaries from the MIMIC-IV database, with clinician ground-truth labels, we benchmark 11 LLMs. Model-assessed mean documentation completeness ranges from 54.9% to 74.2%, and the best-performing models achieve a Cohen's kappa values around 0.5 against clinician labels, indicating moderate agreement. All models struggle to identify ambiguous documentation (Unclear), highlighting a key gap in current automated auditing. This work provides a clinician-validated benchmark and zero-shot baselines for systematic quality improvement in clinical documentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。