arXiv:2608.16725cs.CVcs.AI2026-08

用无监督方法自动检测多中心乳腺MRI数据异常,提升医疗AI可靠性。

Unsupervised Anomaly Detection for Image Dataset Quality Assurance in Multi-Center Breast MRI

论文配图:Unsupervised Anomaly Detection for Image Dataset Quality Assurance in Multi-Center Breast MRI
图 1 · 摘自论文原文
  • 基于视觉感知构建放射影像异常分类体系,支持细粒度分析。
  • 3D重建法在跨机构泛化上表现最佳(AUROC 0.936),投影法整体最优(0.954)。
  • 植入物和乳房切除术后影像仍是未解难题,方法需领域适配。

损坏、不一致或异常数据会隐性威胁医疗AI的安全与可靠性。尽管监管层日益重视高风险医疗AI的数据集质量保障(QA),但可扩展的自动化检测仍不成熟。本文采用无监督异常检测(AD)与分布外(OOD)检测,构建多中心动态对比增强乳腺MRI的数据集QA机制。基于六个公开数据集,创建包含17种真实场景异常类型的可控基准,涵盖协议违规、处理错误、解剖区域误判等。提出基于人类视觉感知的放射影像异常分类体系,支持对AD失败模式的细粒度分析。基准包含近-、中远-、远距离OOD样本,以及分布内与外部正常数据。评估四种方法:一种结合领域特定特征提取器与新型位置编码的投影法;一种扩展至全3D体数据并引入增强训练目标的重构法;以及两种未经修改的混合式OOD检测法。中远及远距离OOD样本检测可靠,而近距离OOD样本与未见机构的外部正常数据暴露方法差异。3D重构法在检测性能(AUROC: 0.936)与跨机构泛化间取得最佳平衡;结合位置编码的投影法总体表现最优(AUROC: 0.954)。两种混合方法均出现关键失败模式,表明仅在单一模态或解剖结构验证的方法,若无领域适配则难以泛化。植入物与乳房切除术后影像对所有方法仍是挑战。研究为医疗AI流程中可扩展的无监督质检提供基础与实用指导。

原文摘要 · Abstract (English)

Corrupted, inconsistent, or anomalous data silently threatens the safety and reliability of medical AI. Despite growing regulatory recognition of dataset quality assurance (QA) for high-risk medical AI, scalable automated detection remains underdeveloped. We employ unsupervised anomaly detection (AD) and out-of-distribution (OOD) detection as an automated dataset QA mechanism for multi-center dynamic contrast-enhanced breast MRI. We build a controlled AD benchmark of 17 realistic QA-relevant anomaly types from six public datasets (protocol violations, processing errors, incorrect anatomical regions) and propose a taxonomy of radiological image anomalies based on human visual perception, enabling fine-grained analysis of AD failure modes. The benchmark includes near-, medium-far-, far-OOD samples, as well as in-distribution and external normal data. Four methods are evaluated: a projection-based method extended with a domain-specific feature extractor and a novel positional encoding, a reconstruction-based approach extended to full 3D volumes with an augmented training objective, and two unmodified hybrid OOD detection methods. Medium-far- and far-OOD samples are detected reliably, whereas near-OOD samples and external normal data from unseen institutions expose method-specific differences. The 3D reconstruction-based approach best balances detection performance (AUROC: 0.936) and generalization to unseen institutions. The projection-based method with positional encoding achieves the highest overall detection performance (AUROC: 0.954). Both hybrid methods exhibit critical failure modes, confirming that methods validated for one modality or anatomy may not generalize without domain-specific adaptation. Implants and mastectomies remain an open challenge for all methods. Our results establish a foundation and practical guidance on scalable unsupervised QA in medical AI pipelines.

异常检测医疗AI乳腺MRI数据质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。