arXiv:2502.10562cs.CVcs.LG2025-02被引 3

检测乳腺癌AI在不同人群中的性能偏差并实时预警

Detecting and Monitoring Bias for Subgroups in Breast Cancer Detection AI

  • 按六类属性分组评估模型表现,识别出表现欠佳的子群体
  • 发现部分子群体准确率显著下降,存在潜在算法偏见
  • 提出动态监控机制,支持及时干预与持续优化

自动化乳腺钼靶筛查在早期乳腺癌检测中具有重要作用。然而,当前基于某些训练数据集开发的机器学习模型,在真实部署环境中可能出现性能下降和偏差问题。本文分析了高性能AI模型在两个乳腺影像数据集——埃默里乳腺影像数据集(EMBED)和RSNA 2022挑战赛数据集——上的表现,特别考察了模型在六个属性定义的子群体中的分类性能,使用多种评估指标检测潜在偏差。分析发现部分子群体表现明显落后,凸显了对这些群体进行持续监测的必要性。为此,我们采用一种性能漂移检测方法,一旦发现性能下降即触发警报,从而支持及时干预。该方法不仅提供性能追踪工具,也保障了AI模型在多样化人群中的长期有效性。

原文摘要 · Abstract (English)

Automated mammography screening plays an important role in early breast cancer detection. However, current machine learning models, developed on some training datasets, may exhibit performance degradation and bias when deployed in real-world settings. In this paper, we analyze the performance of high-performing AI models on two mammography datasets-the Emory Breast Imaging Dataset (EMBED) and the RSNA 2022 challenge dataset. Specifically, we evaluate how these models perform across different subgroups, defined by six attributes, to detect potential biases using a range of classification metrics. Our analysis identifies certain subgroups that demonstrate notable underperformance, highlighting the need for ongoing monitoring of these subgroups' performance. To address this, we adopt a monitoring method designed to detect performance drifts over time. Upon identifying a drift, this method issues an alert, which can enable timely interventions. This approach not only provides a tool for tracking the performance but also helps ensure that AI models continue to perform effectively across diverse populations.

乳腺癌检测算法偏见AI监控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。