用临床预测任务评估大模型病历摘要的实用价值
DistillNote: Toward a Functional Evaluation Framework of LLM-Generated Clinical Note Summaries
- 将摘要用于心衰预测任务,量化诊断信息保留程度
- 压缩至原大小1/20仍保持97%预测性能(AUROC 0.92)
- 适合医疗AI落地评估,可推广至其他疾病和任务
大型语言模型(LLMs)越来越多地用于生成临床病历摘要,但其在保留关键诊断信息方面的能力尚未充分研究,可能对患者安全造成风险。本研究提出DistillNote评估框架,通过将生成的摘要应用于复杂临床预测任务,直接量化其中保留的预测信号。我们基于MIMIC-IV病历数据,以不同压缩率生成超过19.2万条摘要:标准、分节压缩、逐节压缩。选取心衰诊断作为预测任务,因其需整合广泛临床信号。在原始病历与摘要上分别微调模型,并使用AUROC指标比较诊断性能。结果表明,即使压缩至原大小约1/20,模型在摘要上训练仍达到AUROC 0.92,仅比原始病历基准(AUROC 0.94)低3%,即保留97%预测信号。该功能评估提供了衡量医疗摘要质量的新视角,强调临床实用性。DistillNote是一种可扩展的任务驱动评估方法,首次系统揭示了临床摘要压缩与性能之间的权衡关系。框架可适配其他预测任务和临床领域,助力真实医疗场景中部署大模型摘要系统的数据驱动决策。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used to generate summaries from clinical notes. However, their ability to preserve essential diagnostic information remains underexplored, which could lead to serious risks for patient care. This study introduces DistillNote, an evaluation framework for LLM summaries that targets their functional utility by applying the generated summary downstream in a complex clinical prediction task, explicitly quantifying how much prediction signal is retained. We generated over 192,000 LLM summaries from MIMIC-IV clinical notes with increasing compression rates: standard, section-wise, and distilled section-wise. Heart failure diagnosis was chosen as the prediction task, as it requires integrating a wide range of clinical signals. LLMs were fine-tuned on both the original notes and their summaries, and their diagnostic performance was compared using the AUROC metric. We contrasted DistillNote's results with evaluations from LLM-as-judge and clinicians, assessing consistency across different evaluation methods. Summaries generated by LLMs maintained a strong level of heart failure diagnostic signal despite substantial compression. Models trained on the most condensed summaries (about 20 times smaller) achieved an AUROC of 0.92, compared to 0.94 with the original note baseline (97 percent retention). Functional evaluation provided a new lens for medical summary assessment, emphasizing clinical utility as a key dimension of quality. DistillNote introduces a new scalable, task-based method for assessing the functional utility of LLM-generated clinical summaries. Our results detail compression-to-performance tradeoffs from LLM clinical summarization for the first time. The framework is designed to be adaptable to other prediction tasks and clinical domains, aiding data-driven decisions about deploying LLM summarizers in real-world healthcare settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。