提出自动监测与数据融合框架,让医疗AI长期稳定运行。
Robust by Design: A Continuous Monitoring and Data Integration Framework for Medical AI
- 用多指标分析+不确定性筛选,自动判断何时更新数据。
- 新数据仅在相似且低不确定时加入,模型性能下降<5%。
- 适合需持续学习的医疗影像场景,防数据漂移与遗忘。
自适应医疗AI模型常因数据漂移导致性能下降。我们提出一种自主的连续监控与数据融合框架,以维持长期鲁棒性。聚焦于肾小球病理图像分类(增生性与非增生性狼疮性肾炎),采用三阶段方法:通过多指标特征分析和基于蒙特卡洛丢弃的不确定性门控,判断是否应使用新数据进行再训练。仅将与训练分布统计相似(欧氏、余弦、马氏距离)且预测熵低的新图像纳入。模型在严格性能保护下增量再训练(任意指标退化不超过5%)。在多中心数据集上使用ResNet18集成模型实验表明,该框架有效防止性能下降:新增数据后AUC保持约0.92,准确率约89%,无显著变化。该方法应对数据偏移,避免灾难性遗忘,实现医疗影像AI的持续学习。
原文摘要 · Abstract (English)
Adaptive medical AI models often face performance drops in dynamic clinical environments due to data drift. We propose an autonomous continuous monitoring and data integration framework that maintains robust performance over time. Focusing on glomerular pathology image classification (proliferative vs. non-proliferative lupus nephritis), our three-stage method uses multi-metric feature analysis and Monte Carlo dropout-based uncertainty gating to decide when to retrain on new data. Only images statistically similar to the training distribution (via Euclidean, cosine, Mahalanobis metrics) and with low predictive entropy are integrated. The model is then incrementally retrained with these images under strict performance safeguards (no metric degradation >5%). In experiments with a ResNet18 ensemble on a multi-center dataset, the framework prevents performance degradation: new images were added without significant change in AUC (~0.92) or accuracy (~89%). This approach addresses data shift and avoids catastrophic forgetting, enabling sustained learning in medical imaging AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。