arXiv:2506.17442cs.AIcs.ET2025-06综述被引 20

医疗AI易随时间失效,这篇综述系统梳理了检测与修复方法。

Keeping Medical AI Healthy and Trustworthy: A Review of Detection and Correction Methods for System Degradation

  • 从数据与模型双角度分析性能退化的根源
  • 提出检测数据漂移、模型退化及根因分析的主流技术
  • 适合关注医疗AI长期可靠性的研究者与临床工程师

人工智能在现代医疗中日益普及,为临床决策提供有力支持。然而,在真实场景中,受数据分布变化、患者特征演变、临床指南更新及数据质量波动等因素影响,AI系统可能随时间出现性能退化,威胁模型可靠性,增加误判风险与不良后果。本文从前瞻性视角出发,系统探讨医疗AI系统的“健康”监测与维护机制,强调持续性能监控、早期退化检测与有效自修正的重要性。首先回顾数据与模型层面导致性能下降的常见原因;随后总结检测数据漂移与模型漂移的关键技术,并深入分析根因定位方法;进一步综述从模型重训练到测试时自适应等各类修正策略。涵盖传统机器学习模型与前沿大语言模型(LLMs),对比其优势与局限。最后讨论现存技术挑战并提出未来研究方向。本工作旨在推动可信赖、鲁棒的医疗AI系统发展,支持其在动态临床环境中安全、长期部署。

原文摘要 · Abstract (English)

Artificial intelligence (AI) is increasingly integrated into modern healthcare, offering powerful support for clinical decision-making. However, in real-world settings, AI systems may experience performance degradation over time, due to factors such as shifting data distributions, changes in patient characteristics, evolving clinical protocols, and variations in data quality. These factors can compromise model reliability, posing safety concerns and increasing the likelihood of inaccurate predictions or adverse outcomes. This review presents a forward-looking perspective on monitoring and maintaining the "health" of AI systems in healthcare. We highlight the urgent need for continuous performance monitoring, early degradation detection, and effective self-correction mechanisms. The paper begins by reviewing common causes of performance degradation at both data and model levels. We then summarize key techniques for detecting data and model drift, followed by an in-depth look at root cause analysis. Correction strategies are further reviewed, ranging from model retraining to test-time adaptation. Our survey spans both traditional machine learning models and state-of-the-art large language models (LLMs), offering insights into their strengths and limitations. Finally, we discuss ongoing technical challenges and propose future research directions. This work aims to guide the development of reliable, robust medical AI systems capable of sustaining safe, long-term deployment in dynamic clinical settings.

医疗AI模型退化持续监控可信计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。