arXiv:2602.00074cs.CYcs.AI2026-02被引 4

医院自建LLM系统,实现病历自动化处理与持续评估

Adoption and Use of LLMs at an Academic Medical Center

  • 开发可对接多年病历的ChatEHR系统,支持固定流程自动化和交互式使用
  • 1.5年部署7个自动化任务,3个月超1万次使用,平均每次生成含0.73幻觉、1.60错误
  • 不依赖厂商,自主管理,估算首年节省600万美元,适合医疗机构落地应用

大型语言模型(LLMs)虽能辅助临床记录,但独立工具因手动录入导致工作流摩擦。我们构建了ChatEHR系统,使LLM可访问覆盖数年跨度的完整患者病历。该系统支持静态提示与数据组合的自动化任务,以及通过用户界面(UI)在电子健康记录(EHR)中进行交互式使用。由此实现多种场景下的病历处理:如就诊前审查、转诊资格筛查、手术部位感染监测和文档抽取,将LLM能力转化为机构级技术。系统经培训后开放使用,支持持续监控与评估。1.5年内建成7项自动化,1075名用户完成培训,上线首月即产生23,000次会话。自动化设计强调模型无偏好性,并可接入多类数据,以匹配具体临床或行政需求。基于基准的评估不足以监控与评价UI表现,需发展新方法。摘要生成是最常见任务,平均每生成一次存在0.73次幻觉和1.60次不准确。成本节约、时间节省及收入增长需价值评估框架来优先排序并量化影响。初步估算首年节省约600万美元,未计入更优诊疗带来的收益。这种‘内部自建’策略为医疗系统提供自主权,建立不依赖供应商、内部可控的LLM平台。

原文摘要 · Abstract (English)

While large language models (LLMs) can support clinical documentation needs, standalone tools struggle with "workflow friction" from manual data entry. We developed ChatEHR, a system that enables the use of LLMs with the entire patient timeline spanning several years. ChatEHR enables automations - which are static combinations of prompts and data that perform a fixed task - and interactive use in the electronic health record (EHR) via a user interface (UI). The resulting ability to sift through patient medical records for diverse use-cases such as pre-visit chart review, screening for transfer eligibility, monitoring for surgical site infections, and chart abstraction, redefines LLM use as an institutional capability. This system, accessible after user-training, enables continuous monitoring and evaluation of LLM use. In 1.5 years, we built 7 automations and 1075 users have trained to become routine users of the UI, engaging in 23,000 sessions in the first 3 months of launch. For automations, being model-agnostic and accessing multiple types of data was essential for matching specific clinical or administrative tasks with the most appropriate LLM. Benchmark-based evaluations proved insufficient for monitoring and evaluation of the UI, requiring new methods to monitor performance. Generation of summaries was the most frequent task in the UI, with an estimated 0.73 hallucinations and 1.60 inaccuracies per generation. The resulting mix of cost savings, time savings, and revenue growth required a value assessment framework to prioritize work as well as quantify the impact of using LLMs. Initial estimates are $6M savings in the first year of use, without quantifying the benefit of the better care offered. Such a "build-from-within" strategy provides an opportunity for health systems to maintain agency via a vendor-agnostic, internally governed LLM platform.

医疗AILLM应用电子病历自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。