无需标签即可估算医疗影像模型在分布偏移下的真实性能。
Label-free estimation of clinically relevant performance metrics under distribution shifts
- 直接估计混淆矩阵,而非仅预测准确率。
- 在真实与模拟分布偏移下,对胸部X光数据表现可靠。
- 揭示现有方法在临床部署中的潜在失效模式,适合医疗AI监管者。
临床图像分类模型的安全部署需要持续性能监控。然而,目标数据集通常缺乏真实标签,导致无法直接评估实际性能。现有先进方法通过利用置信度分数估计目标准确率,但主要关注准确率,且很少在存在严重类别不平衡和数据分布偏移的临床场景中验证。本文贡献有二:一是提出现有性能预测方法的推广,可直接估计完整的混淆矩阵;二是基于真实世界分布偏移及模拟的协变量与先验分布偏移,在胸部X光数据上进行基准测试。所提混淆矩阵估计方法在分布偏移下能可靠预测临床相关的计数指标。然而,模拟偏移实验暴露了当前性能估计技术的重要失效模式,提示在实现医疗AI模型上市后监测时,需更深入理解真实部署环境。
原文摘要 · Abstract (English)
Performance monitoring is essential for safe clinical deployment of image classification models. However, because ground-truth labels are typically unavailable in the target dataset, direct assessment of real-world model performance is infeasible. State-of-the-art performance estimation methods address this by leveraging confidence scores to estimate the target accuracy. Despite being a promising direction, the established methods mainly estimate the model's accuracy and are rarely evaluated in a clinical domain, where strong class imbalances and dataset shifts are common. Our contributions are twofold: First, we introduce generalisations of existing performance prediction methods that directly estimate the full confusion matrix. Then, we benchmark their performance on chest x-ray data in real-world distribution shifts as well as simulated covariate and prevalence shifts. The proposed confusion matrix estimation methods reliably predicted clinically relevant counting metrics on medical images under distribution shifts. However, our simulated shift scenarios exposed important failure modes of current performance estimation techniques, calling for a better understanding of real-world deployment contexts when implementing these performance monitoring techniques for postmarket surveillance of medical AI models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。