arXiv:2604.23954cs.AI2026-04

更新糖尿病模型可能引发预测不稳、随意和不公平,需持续监控。

An empirical evaluation of the risks of AI model updates using clinical data: stability, arbitrariness, and fairness

  • 用4个糖尿病数据集测试不同更新策略对模型影响
  • 更新后超10%病例预测变化,公平性下降且误差分布失衡
  • 提出多维监测框架,适合临床AI系统开发者使用

人工智能与机器学习模型在临床决策支持中日益普及。但当训练数据因人口结构、环境或患者行为变化而过时时,模型性能可能显著下降。尽管定期更新模型是必要的,但更新本身也可能引入新风险。本研究基于四个美国儿童1型糖尿病数据集(共496名20岁以下参与者,约11,300周观测值),利用高分辨率连续血糖监测数据,以预测严重高血糖事件为案例,评估不同模型更新策略对模型稳定性、预测任意性和子群体公平性的影响。结果表明,重训后大量病例的预测发生改变,预测随机性上升,跨人群误差分布不平衡加剧。本文提出多维度持续监控机制,强调其对构建可信临床决策支持系统的重要性。

原文摘要 · Abstract (English)

Artificial Intelligence (AI) and Machine Learning (ML) models used in clinical settings are increasingly deployed to support clinical decision-making. However, when training data become stale due to changes in demographics, environment, or patient behaviors, model performance can degrade substantially. While updating models with new training data is necessary, such updates may also introduce new risks. We evaluated the proposed monitoring framework on four publicly available U.S.-based Type 1 Diabetes datasets containing high-resolution continuous glucose monitoring (CGM) data, comprising approximately 11,300 weekly observations from 496 participants younger than 20 years. All datasets included structured sociodemographic information. Using the prediction of severe hyperglycemia events in children with Type 1 Diabetes as a case study, we examine how different model update strategies can adversely affect model stability by causing predictions to change for a large number of cases after retraining, increase prediction arbitrariness, and worsen subgroup fairness and the balance of error rates across populations. We propose multiple dimensions for continuous monitoring to detect these issues and argue that such monitoring is essential for the development of trustworthy clinical decision support systems.

AI医疗模型更新公平性糖尿病

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。