CNN在癌病理图像分析中可能受隐性偏见影响,导致评估结果不可靠。
Unmasking Biases and Reliability Concerns in Convolutional Neural Networks Analysis of Cancer Pathology Images
- 用背景裁片替代真实病灶,检验模型是否依赖非医学信息
- 部分模型在无临床信息数据上仍达93%准确率,暴露严重偏差
- 提醒研究者警惕基准数据集的误导性,尤其关注模型敏感度差异
卷积神经网络在从影像中识别癌症类型方面表现出良好效果。然而,其内部运作机制不透明,限制了对其性能的深入理解,仅能依赖经验评估。本文针对癌症病理分析中常用的评估范式进行验证,分析了13个广泛使用的癌症基准数据集,采用四种常见CNN架构,涵盖黑色素瘤、癌、结直肠癌和肺癌等类型。将原始图像的背景区域裁剪后生成不含临床信息的数据集,对比模型在这些无医学内容数据上的分类准确率。根据零假设,此类数据应仅产生随机水平的准确率。结果表明,部分模型在裁剪数据上仍达到高达93%的准确率,说明其可能过度依赖图像背景中的非医学特征。这揭示了某些CNN架构对偏差更敏感,当前主流机器学习评估方法在癌症病理领域可能导致不可靠结论。这些偏差极难察觉,可能误导研究者误判模型有效性。
原文摘要 · Abstract (English)
Convolutional Neural Networks have shown promising effectiveness in identifying different types of cancer from radiographs. However, the opaque nature of CNNs makes it difficult to fully understand the way they operate, limiting their assessment to empirical evaluation. Here we study the soundness of the standard practices by which CNNs are evaluated for the purpose of cancer pathology. Thirteen highly used cancer benchmark datasets were analyzed, using four common CNN architectures and different types of cancer, such as melanoma, carcinoma, colorectal cancer, and lung cancer. We compared the accuracy of each model with that of datasets made of cropped segments from the background of the original images that do not contain clinically relevant content. Because the rendered datasets contain no clinical information, the null hypothesis is that the CNNs should provide mere chance-based accuracy when classifying these datasets. The results show that the CNN models provided high accuracy when using the cropped segments, sometimes as high as 93\%, even though they lacked biomedical information. These results show that some CNN architectures are more sensitive to bias than others. The analysis shows that the common practices of machine learning evaluation might lead to unreliable results when applied to cancer pathology. These biases are very difficult to identify, and might mislead researchers as they use available benchmark datasets to test the efficacy of CNN methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。