构建可落地的临床AI持续治理框架,提升系统可靠性与医生体验
End-to-End Evaluation and Governance of an EHR-Embedded AI Agent for Clinicians

- 设计多通道治理框架,融合评分验证、实时反馈与性能监控
- 7个版本迭代后平均评分从84%提升至95%,错误报告减少超60%
- 适合医疗AI部署团队、临床研究者及医院信息化管理者参考
临床AI系统不仅需要一次性评估,更需持续治理——即在部署过程中持续监控、评估、迭代与再评估。本文提出一个端到端治理框架,整合评分验证、实时部署反馈、技术性能监控与成本追踪,并通过受控实验机制在发布前验证变更。该框架应用于嵌入电子病历(EHR)的Hyperscribe系统,该系统将环境音频转为结构化病历更新。20名临床医生针对823例病例创建了1,646份经验证的评分标准。七个Hyperscribe版本通过受控实验评估,平均得分从84%提升至95%。三个月内收集107条实时反馈显示,错误报告占比从79%降至30%,积极反馈从14%升至45%,表明工程干预有效修复问题。每段音频平均处理时间为8.1秒,经重试机制后有效完成率达99.6%。结果表明,持续的多通道治理在临床AI部署中既可行又高效。
原文摘要 · Abstract (English)
Clinical AI systems require not just point-in-time evaluation but continuous governance: the ongoing practice of monitoring, evaluating, iterating, and re-evaluating performance throughout deployment. We present an end-to-end framework of governance that integrates rubric validation, live deployment feedback, technical performance monitoring, and cost tracking, with controlled experimentation gating system changes before deployment. Applied to Hyperscribe, an EHR-embedded agent that converts ambient audio into structured chart updates, twenty clinicians authored 1,646 validated rubrics across 823 cases. Seven Hyperscribe versions were evaluated through controlled experiments, with median scores improving from 84% to 95%. Analysis of 107 live feedback entries over three months showed feedback composition shifting from 79% error reports and 14% positive observations to 30% errors and 45% positive observations as engineering interventions resolved failures. Median processing time per audio segment was 8.1 seconds with a 99.6% effective completion rate after retry mechanisms absorbed transient model errors. These results demonstrate that continuous, multi-channel governance of deployed clinical AI is both achievable and effective.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。