arXiv:2506.05701cs.LG2025-06NeurIPS被引 15

为临床AI部署后监控建立统计有效的测试框架,提升系统可靠性与安全性。

Statistically Valid Post-Deployment Monitoring Should Be Standard for AI-Based Digital Health

  • 将数据与模型性能变化视为独立的统计假设检验问题。
  • 仅9%的FDA注册医疗AI工具有部署后监控计划,现有方法多为手动且被动。
  • 适合关注临床AI安全、监管合规与可复现性研究的团队。

本文主张临床AI的部署后监控严重不足,应建立基于标签高效和统计有效的测试框架,作为保障真实世界部署中可靠性和安全性的原则基础。近期审查发现,仅有9%的FDA注册医疗AI工具包含部署后监测计划。当前监控方法多为人工、零星且被动,难以适应临床模型运行的动态环境。本文认为,部署后监控应建立在提供明确误差率保证(如Ⅰ/Ⅱ类错误)、支持预设假设下的正式推断及可复现性的统计框架之上,契合监管要求。具体而言,数据分布变化与模型性能退化应分别建模为统计假设检验问题。以统计严谨性为基础,可为临床AI系统的持续可靠性提供科学支撑,并开启关于检测、归因与缓解真实场景中模型失效的技术研究新方向。

原文摘要 · Abstract (English)

This position paper argues that post-deployment monitoring in clinical AI is underdeveloped and proposes statistically valid and label-efficient testing frameworks as a principled foundation for ensuring reliability and safety in real-world deployment. A recent review found that only 9% of FDA-registered AI-based healthcare tools include a post-deployment surveillance plan. Existing monitoring approaches are often manual, sporadic, and reactive, making them ill-suited for the dynamic environments in which clinical models operate. We contend that post-deployment monitoring should be grounded in label-efficient and statistically valid testing frameworks, offering a principled alternative to current practices. We use the term "statistically valid" to refer to methods that provide explicit guarantees on error rates (e.g., Type I/II error), enable formal inference under pre-defined assumptions, and support reproducibility--features that align with regulatory requirements. Specifically, we propose that the detection of changes in the data and model performance degradation should be framed as distinct statistical hypothesis testing problems. Grounding monitoring in statistical rigor ensures a reproducible and scientifically sound basis for maintaining the reliability of clinical AI systems. Importantly, it also opens new research directions for the technical community--spanning theory, methods, and tools for statistically principled detection, attribution, and mitigation of post-deployment model failures in real-world settings.

AI医疗部署监控统计验证临床安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。