自反思智能体可自动检测临床AI模型的失效风险并修复性能下降
An autonomous agent for auditing and improving the reliability of clinical AI models
- 构建多智能体系统,模拟真实医疗场景下的数据分布偏移
- 在3个真实临床场景中识别出模型性能下降,修复后恢复15%-25%损失
- 可在消费级硬件上10分钟内完成审计,成本低于0.5美元
临床AI模型部署面临严峻挑战:在基准测试中表现接近专家水平的模型,在面对真实世界医学影像中的设备差异、光照变化或人群特征变化时可能突然失效。当前可靠性审计依赖定制化流程,耗时且缺乏可解释性工具。本文提出ModelAuditor,一个能与用户对话、选择任务相关指标,并模拟临床相关的分布偏移的自反思智能体。它生成可解释报告,预测部署中性能下降程度,分析具体失效模式,定位根本原因并提出缓解策略。在三个真实场景——组织病理学跨机构差异、皮肤科人群分布变化、胸部X光设备异质性中,ModelAuditor成功识别了先进模型(如SIIM-ISIC melanoma classifier)的上下文特异性失效模式。其推荐策略使因分布偏移导致的15%-25%性能损失得以恢复,显著优于基线模型和现有增强方法。该系统采用多智能体架构,可在消费级硬件上于10分钟内完成审计,单次成本低于0.5美元。
原文摘要 · Abstract (English)
The deployment of AI models in clinical practice faces a critical challenge: models achieving expert-level performance on benchmarks can fail catastrophically when confronted with real-world variations in medical imaging. Minor shifts in scanner hardware, lighting or demographics can erode accuracy, but currently reliability auditing to identify such catastrophic failure cases before deployment is a bespoke and time-consuming process. Practitioners lack accessible and interpretable tools to expose and repair hidden failure modes. Here we introduce ModelAuditor, a self-reflective agent that converses with users, selects task-specific metrics, and simulates context-dependent, clinically relevant distribution shifts. ModelAuditor then generates interpretable reports explaining how much performance likely degrades during deployment, discussing specific likely failure modes and identifying root causes and mitigation strategies. Our comprehensive evaluation across three real-world clinical scenarios - inter-institutional variation in histopathology, demographic shifts in dermatology, and equipment heterogeneity in chest radiography - demonstrates that ModelAuditor is able correctly identify context-specific failure modes of state-of-the-art models such as the established SIIM-ISIC melanoma classifier. Its targeted recommendations recover 15-25% of performance lost under real-world distribution shift, substantially outperforming both baseline models and state-of-the-art augmentation methods. These improvements are achieved through a multi-agent architecture and execute on consumer hardware in under 10 minutes, costing less than US$0.50 per audit.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。