用病历元数据去除冗余文本,让大模型更高效地处理临床记录。
Clinical Note Bloat Reduction for Efficient LLM Use

- 利用电子病历元数据识别模板化和复制内容,自动去重。
- 在530万份病历中减少47.3%的文本量,不影响诊断预测性能。
- 适合希望降低LLM部署成本的医疗机构与系统开发者。
医疗系统正快速部署大型语言模型(LLMs)用于临床决策支持,但现代病历书写依赖模板、复制粘贴和自动生成字段,导致大量重复文本(“病历膨胀”),削弱了临床信号并显著增加计算成本。我们提出TRACE——一种可扩展的预处理流程,通过利用电子病历(EHR)的归属元数据识别模板化和复制内容;当元数据不可用时,则采用基于频率的去重策略。我们在涵盖肝移植、产科和住院护理的四个真实临床队列(共530万份病历)中评估TRACE,采用盲法医生评审和下游建模任务。结果表明,TRACE在移除47.3%病历文本的同时,保持了信息提取与临床结局预测的性能。在一所大型学术医学中心,该优化预计每年可节省950万美元的LLM推理成本(按每例就诊一次查询估算)。研究证明,未被充分利用的EHR元数据能显著提升基于LLM的临床系统部署的可扩展性与经济性。
原文摘要 · Abstract (English)
Health systems are rapidly deploying large language models (LLMs) that use clinical notes for clinical decision support applications. However, modern documentation practices rely heavily on templates, copy--paste shortcuts, and auto-populated fields, producing extensive duplicated text (``note bloat'') that dilutes clinically meaningful signal and substantially increases the computational cost of LLM use. We introduce TRACE, a scalable preprocessing pipeline that removes note bloat by leveraging EHR attribution metadata to identify templated and copied content and applying frequency-based deduplication when metadata are unavailable. We evaluated TRACE across four real--world clinical cohorts spanning liver transplantation, obstetrics, and inpatient care (5.3 million notes) using blinded physician review and downstream modeling tasks. TRACE removed 47.3% of chart text while preserving performance for information extraction and clinical outcome prediction. At a large academic medical center, this reduction corresponds to an estimated $9.5 million annual decrease in LLM inference costs assuming one query per encounter. These findings show how underutilized EHR metadata can enable more scalable and cost-efficient deployment of LLM-based clinical systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。