arXiv:2512.01534cs.CVcs.AI2025-12被引 1

构建脑影像异常检测大规模基准,揭示现有方法的偏差与局限。

Deep Unsupervised Anomaly Detection in Brain Imaging: Large-Scale Benchmarking and Bias Analysis

  • 建立多中心、大样本脑影像无监督异常检测基准,涵盖6种扫描仪和多样人群。
  • 重建类方法(如扩散模型)分割效果最好,但多数算法对小病灶和低对比度病变漏检严重。
  • 发现扫描仪差异、年龄性别等系统性偏差普遍存在,提示需改进公平性与鲁棒性设计。

脑磁共振成像中的深度无监督异常检测有望在无需病灶标注的情况下识别病理异常,但碎片化评估、异质数据集和不一致指标阻碍了临床转化。本文构建了一个大规模多中心基准,训练集包含来自6台扫描仪的2,976例T1和2,972例T2加权健康扫描,年龄6至89岁;验证集使用92例扫描调参并估计无偏阈值;测试集涵盖2,221例T1w和1,262例T2w扫描,覆盖健康及多种临床队列。所有算法的基于骰子系数的分割性能在0.03至0.65之间,差异显著。系统评估显示,重建类方法(尤其是扩散启发式方法)表现最优,特征类方法在分布外迁移中更稳健。然而,多数算法存在系统性偏差:小病灶和低对比度病灶漏检率高,假阳性随年龄和性别变化。增加健康数据仅带来小幅提升,表明当前框架瓶颈在于算法而非数据。该基准为未来研究提供透明基础,强调临床转化关键方向:原生图像预训练、合理偏离度量、公平性建模与鲁棒域适应。

原文摘要 · Abstract (English)

Deep unsupervised anomaly detection in brain magnetic resonance imaging offers a promising route to identify pathological deviations without requiring lesion-specific annotations. Yet, fragmented evaluations, heterogeneous datasets, and inconsistent metrics have hindered progress toward clinical translation. Here, we present a large-scale, multi-center benchmark of deep unsupervised anomaly detection for brain imaging. The training cohort comprised 2,976 T1 and 2,972 T2-weighted scans from healthy individuals across six scanners, with ages ranging from 6 to 89 years. Validation used 92 scans to tune hyperparameters and estimate unbiased thresholds. Testing encompassed 2,221 T1w and 1,262 T2w scans spanning healthy datasets and diverse clinical cohorts. Across all algorithms, the Dice-based segmentation performance varied between 0.03 and 0.65, indicating substantial variability. To assess robustness, we systematically evaluated the impact of different scanners, lesion types and sizes, as well as demographics (age, sex). Reconstruction-based methods, particularly diffusion-inspired approaches, achieved the strongest lesion segmentation performance, while feature-based methods showed greater robustness under distributional shifts. However, systematic biases, such as scanner-related effects, were observed for the majority of algorithms, including that small and low-contrast lesions were missed more often, and that false positives varied with age and sex. Increasing healthy training data yields only modest gains, underscoring that current unsupervised anomaly detection frameworks are limited algorithmically rather than by data availability. Our benchmark establishes a transparent foundation for future research and highlights priorities for clinical translation, including image native pretraining, principled deviation measures, fairness-aware modeling, and robust domain adaptation.

脑影像异常检测无监督学习基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。