arXiv:2509.26153cs.AI2025-09

如何把AI医生落地到医院,关键不在算法,在流程与协作。

A Field Guide to Deploying AI Agents in Clinical Practice

  • 用真实临床数据训练的AI助手,自动筛查免疫治疗副作用。
  • 80%精力花在数据对接、合规审查等非算法工作上。
  • 适合医院信息科、临床工程师和医疗AI项目负责人读。

将大语言模型融入代理式工作流,在医疗领域潜力巨大,但实际落地仍存巨大鸿沟。本文基于在麻省总医院布里格姆系统部署‘irAE-Agent’(用于从临床笔记中自动识别免疫相关不良事件)的经验,以及对21位临床医生、工程师和信息学领导的结构化访谈,提出一份面向实践者的操作手册。分析显示,临床AI开发中不足20%的时间用于提示工程与模型构建,超过80%耗费在社会技术层面的实施工作。我们提炼出五大核心挑战:数据集成、模型验证、经济价值保障、系统漂移管理与治理机制。通过提供针对每项挑战的可行动方案,本手册将关注点从算法开发转向支撑系统落地所需的基础设施与实施工作,旨在跨越‘死亡谷’,推动生成式AI从试点走向日常临床应用。

原文摘要 · Abstract (English)

Large language models (LLMs) integrated into agent-driven workflows hold immense promise for healthcare, yet a significant gap exists between their potential and practical implementation within clinical settings. To address this, we present a practitioner-oriented field manual for deploying generative agents that use electronic health record (EHR) data. This guide is informed by our experience deploying the "irAE-Agent", an automated system to detect immune-related adverse events from clinical notes at Mass General Brigham, and by structured interviews with 21 clinicians, engineers, and informatics leaders involved in the project. Our analysis reveals a critical misalignment in clinical AI development: less than 20% of our effort was dedicated to prompt engineering and model development, while over 80% was consumed by the sociotechnical work of implementation. We distill this effort into five "heavy lifts": data integration, model validation, ensuring economic value, managing system drift, and governance. By providing actionable solutions for each of these challenges, this field manual shifts the focus from algorithmic development to the essential infrastructure and implementation work required to bridge the "valley of death" and successfully translate generative AI from pilot projects into routine clinical care.

AI医疗临床落地生成式AI电子病历

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。