科学影像中深度学习会失效,因数据特性与模型偏见不匹配。
Anatomy of a failure: When, how, and why deep vision fails in scientific domains

- 用红外光谱数据对比病理图像,发现模型性能反而不如普通图像。
- 模型因数据先验与深度学习的简单性偏好冲突,退化为一维预测。
- 提醒科研人员警惕通用AI在科学领域的安全风险,需定制算法。
深度学习在科学成像中的应用日益广泛,但其在日常RGB图像上的成功未必适用于科学影像。科学图像像素蕴含数千通道的精确理化属性,而深度学习在人类感知任务中表现良好,但在科学成像中效果存疑。本文通过比较染色组织的RGB图像与红外(IR)成像的定量生物化学信号,发现尽管红外数据信息量更丰富,基于其训练的深度学习模型反而表现更差。研究揭示,这是由于红外数据的先验分布与深度学习固有的简单性偏差不兼容,导致模型坍缩至一维预测,严重浪费表征能力。这一失败具有灾难性,引发人工智能安全性问题,并削弱科学模态优势。值得注意的是,即使采用当前最先进的鲁棒化策略,问题依然存在——这些策略主要针对RGB图像设计和验证,同样存在先验偏差不匹配。本工作建立理解通用深度学习在科学领域局限性的框架,倡导研究模态特异性失败模式,以指导安全、专用的人工智能算法开发。
原文摘要 · Abstract (English)
Mirroring its ubiquity in popular media and all human activities, the use of deep learning (DL) is rapidly growing in scientific imaging modalities. However, unlike everyday RGB pictures, pixels encode precise physicochemical properties in scientific imaging across potentially thousands of channels. While DL is well validated on human-centric RGB perceptual tasks, its effectiveness for scientific imaging remains uncertain. Here, we show that the naive application of DL frameworks to scientific images can lead to critical failures. We evaluate the use of DL for pathology, comparing RGB images of stained tissue with the quantitative and information-rich biochemical signatures of infrared (IR) imaging. Despite this informational advantage, DL models trained on IR data paradoxically underperform. We investigate this discrepancy to find that IR data priors interact poorly with the simplicity bias of DL, causing models to collapse to one-dimensional predictions. This constitutes a catastrophic DL failure because the model's representational capacity remains largely unused, while furthermore raising AI safety concerns and undermining the advantages of such scientific modalities. Notably, this problem persists even with state-of-the-art DL robustification strategies, which are primarily designed and validated for RGB imagery and thus inherit the same prior-bias mismatch. This work establishes a framework for understanding the limitations of generic DL in science and advocates for the study of modality-specific failure modes to guide the development of specialized, safe AI algorithms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。