为医疗AI部署后监控提供可落地的三原则框架,确保安全有效。
Monitoring Deployed AI Systems in Health Care
- 基于系统完整性、性能与影响三原则设计监控机制
- 明确指标、审查频率、责任人及异常时的行动方案
- 适用于传统与生成式AI,兼顾资源有限机构的实操性
医疗AI系统部署后的监控对保障其安全性、质量与持续价值至关重要,有助于指导系统更新、修改或退役的治理决策。为此,我们开发了一套以“系统失效时必须采取行动”为前提的监控框架,已在斯坦福健康中心实际应用。该框架围绕三个互补原则展开:系统完整性、性能与影响。系统完整性监控关注最大化系统可用性,检测运行时错误,并识别周边IT环境变化带来的意外影响;性能监控确保系统在医疗实践和输入数据随时间演变的情况下仍保持准确行为;影响监控评估系统是否持续为临床医生和患者带来实际效益。结合我院已部署AI系统的实例,我们提供了基于上述原则制定监控计划的实用指南,明确应监测的指标、审查时间、责任主体及异常响应措施,涵盖传统与生成式AI。同时讨论了实施挑战,包括资源有限机构的监控成本,以及在目标冲突、成功定义不一的复杂组织中推行数据驱动监控的困难。该框架为医疗机构确保AI长期安全有效提供了可操作的模板。
原文摘要 · Abstract (English)
Post-deployment monitoring of artificial intelligence (AI) systems in health care is essential to ensure their safety, quality, and sustained benefit-and to support governance decisions about which systems to update, modify, or decommission. Motivated by these needs, we developed a framework for monitoring deployed AI systems grounded in the mandate to take specific actions when they fail to behave as intended. This framework, which is now actively used at Stanford Health Care, is organized around three complementary principles: system integrity, performance, and impact. System integrity monitoring focuses on maximizing system uptime, detecting runtime errors, and identifying when changes to the surrounding IT ecosystem have unintended effects. Performance monitoring focuses on maintaining accurate system behavior in the face of changing health care practices (and thus input data) over time. Impact monitoring assesses whether a deployed system continues to have value in the form of benefit to clinicians and patients. Drawing on examples of deployed AI systems at our academic medical center, we provide practical guidance for creating monitoring plans based on these principles that specify which metrics to measure, when those metrics should be reviewed, who is responsible for acting when metrics change, and what concrete follow-up actions should be taken-for both traditional and generative AI. We also discuss challenges to implementing this framework, including the effort and cost of monitoring for health systems with limited resources and the difficulty of incorporating data-driven monitoring practices into complex organizations where conflicting priorities and definitions of success often coexist. This framework offers a practical template and starting point for health systems seeking to ensure that AI deployments remain safe and effective over time.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。