arXiv:2411.07940cs.AIcs.CV2024-11中稿 · MICCAI 2025被引 9

首次实现医学影像数据分布偏移的无监督精准识别,助力AI安全部署。

Automatic dataset shift identification to support safe deployment of medical imaging AI

  • 基于自监督编码器与任务模型输出,构建无监督偏移识别框架。
  • 在三种影像模态上验证,可区分流行率、协变量及混合偏移。
  • 适用于真实世界场景,帮助医生判断AI是否可信可用。

数据分布变化会显著降低临床AI模型性能,导致误诊。现有方法仅能检测偏移存在,但无法识别具体类型,而不同偏移需不同应对策略。本文提出首个针对医学影像的无监督数据分布偏移识别框架,可有效区分流行率偏移(标签分布变化)、协变量偏移(输入特征变化)和混合偏移(两者同时发生)。研究强调自监督编码器对检测细微协变量偏移的重要性,并设计新型检测器融合自监督特征与任务模型输出,提升识别精度。在胸片、数字乳腺摄影和视网膜眼底图像三类模态上,基于五个公开大型数据集,验证了该框架在五种真实世界偏移场景下的有效性。

原文摘要 · Abstract (English)

Shifts in data distribution can substantially harm the performance of clinical AI models and lead to misdiagnosis. Hence, various methods have been developed to detect the presence of such shifts at deployment time. However, the root causes of dataset shifts are diverse, and the choice of shift mitigation strategies is highly dependent on the precise type of shift encountered at test time. As such, detecting test-time dataset shift is not sufficient: precisely identifying which type of shift has occurred is critical. In this work, we propose the first unsupervised dataset shift identification framework for imaging datasets, effectively distinguishing between prevalence shift (caused by a change in the label distribution), covariate shift (caused by a change in input characteristics) and mixed shifts (simultaneous prevalence and covariate shifts). We discuss the importance of self-supervised encoders for detecting subtle covariate shifts and propose a novel shift detector leveraging both self-supervised encoders and task model outputs for improved shift detection. We show the effectiveness of the proposed shift identification framework across three different imaging modalities (chest radiography, digital mammography, and retinal fundus images) on five types of real-world dataset shifts using five large publicly available datasets.

医学影像数据偏移AI安全无监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。