arXiv:2512.18450cs.AIcs.CV2025-12中稿 · MICAD

提出智能体框架,实时检测多中心医疗AI系统中的预测漂移。

Agent-Based Output Drift Detection for Breast Cancer Response Prediction in a Multisite Clinical Decision Support System

  • 每个医疗中心部署智能体,基于输出分布对比识别漂移
  • 自适应方案在真实乳腺癌数据上实现74.3%的漂移检测F1-score
  • 适合多机构部署的临床AI系统,提升模型可靠性

现代临床决策支持系统可同时服务多个独立医学影像机构,但因患者群体、成像设备和采集协议差异,其预测性能可能在不同站点下降。持续监控模型输出是无需真实标签即可识别分布偏移的安全可靠方法。然而,现有方法多依赖集中式聚合预测监控,忽视站点特异性漂移动态。本文提出一种基于智能体的框架,用于检测和评估多中心临床AI系统中的漂移。我们在真实乳腺癌影像数据上模拟多中心环境,为每个站点分配一个漂移监测智能体,通过批处理方式将模型输出与参考分布进行比较。分析了四种参考获取方式(站点专属、全局、仅生产、自适应)及集中式基线方案。结果表明,所有多中心方案均优于集中监控,漂移检测F1-score最高提升10.3%。无站点专属参考时,自适应方案表现最佳,漂移检测F1-score达74.3%,漂移严重性分类达83.7%。结果表明,自适应、站点感知的智能体漂移监控可增强多中心临床决策系统的可靠性。

原文摘要 · Abstract (English)

Modern clinical decision support systems can concurrently serve multiple, independent medical imaging institutions, but their predictive performance may degrade across sites due to variations in patient populations, imaging hardware, and acquisition protocols. Continuous surveillance of predictive model outputs offers a safe and reliable approach for identifying such distributional shifts without ground truth labels. However, most existing methods rely on centralized monitoring of aggregated predictions, overlooking site-specific drift dynamics. We propose an agent-based framework for detecting drift and assessing its severity in multisite clinical AI systems. To evaluate its effectiveness, we simulate a multi-center environment for output-based drift detection, assigning each site a drift monitoring agent that performs batch-wise comparisons of model outputs against a reference distribution. We analyse several multi-center monitoring schemes, that differ in how the reference is obtained (site-specific, global, production-only and adaptive), alongside a centralized baseline. Results on real-world breast cancer imaging data using a pathological complete response prediction model shows that all multi-center schemes outperform centralized monitoring, with F1-score improvements up to 10.3% in drift detection. In the absence of site-specific references, the adaptive scheme performs best, with F1-scores of 74.3% for drift detection and 83.7% for drift severity classification. These findings suggest that adaptive, site-aware agent-based drift monitoring can enhance reliability of multisite clinical decision support systems.

AI医疗漂移检测多中心

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。