MRI Alzheimer病分类模型依赖头骨剥离后的轮廓,而非真实脑组织纹理。
Skull-stripping induces shortcut learning in MRI-based Alzheimer's disease classification
- 通过对比不同预处理方式,发现模型主要依赖头骨剥离产生的脑轮廓
- 即使图像纹理被二值化,分类准确率仍稳定在90%以上
- 揭示了预处理步骤可能引入误导性线索,影响模型可靠性
目的:尽管基于结构磁共振成像(sMRI)的深度神经网络已实现阿尔茨海默病(AD)高精度分类,但其决策依据的图像特征尚不明确。本研究系统评估了T1加权(T1w)灰白质纹理、体积信息及预处理——尤其是头骨剥离——的作用。方法:使用来自ADNI数据库的990例匹配的T1w MRI数据,涵盖阿尔茨海默病患者与认知正常对照。通过头骨剥离和强度二值化等不同预处理策略,分离纹理与形状贡献。在每种配置下训练3D卷积神经网络,并采用精确麦内马尔检验结合离散邦弗朗尼-霍姆校正比较分类性能。利用层间相关性传播、图像相似性度量及相关性图谱的谱聚类分析特征重要性。结果:尽管图像内容差异显著,分类准确率、敏感性和特异性在不同预处理条件下保持稳定。在二值化图像上训练的模型仍维持高性能,表明模型对灰白质纹理依赖极小。相反,体积特征——特别是头骨剥离引入的脑轮廓——始终被模型所依赖。结论:该行为反映了一种捷径学习现象,即预处理伪影成为潜在的非预期提示。由此产生的‘聪明汉斯效应’强调了可解释性工具的重要性,以揭示隐藏偏差,确保医学影像中深度学习的鲁棒性与可信性。
原文摘要 · Abstract (English)
Objectives: High classification accuracy of Alzheimer's disease (AD) from structural MRI has been achieved using deep neural networks, yet the specific image features contributing to these decisions remain unclear. In this study, the contributions of T1-weighted (T1w) gray-white matter texture, volumetric information, and preprocessing -- particularly skull-stripping -- were systematically assessed. Methods: A dataset of 990 matched T1w MRIs from AD patients and cognitively normal controls from the ADNI database were used. Preprocessing was varied through skull-stripping and intensity binarization to isolate texture and shape contributions. A 3D convolutional neural network was trained on each configuration, and classification performance was compared using exact McNemar tests with discrete Bonferroni-Holm correction. Feature relevance was analyzed using Layer-wise Relevance Propagation, image similarity metrics, and spectral clustering of relevance maps. Results: Despite substantial differences in image content, classification accuracy, sensitivity, and specificity remained stable across preprocessing conditions. Models trained on binarized images preserved performance, indicating minimal reliance on gray-white matter texture. Instead, volumetric features -- particularly brain contours introduced through skull-stripping -- were consistently used by the models. Conclusions: This behavior reflects a shortcut learning phenomenon, where preprocessing artifacts act as potentially unintended cues. The resulting Clever Hans effect emphasizes the critical importance of interpretability tools to reveal hidden biases and to ensure robust and trustworthy deep learning in medical imaging.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。