发现医学影像ViT表示不具语义意义,微小变化即致分类失效
Are Vision Transformer Representations Semantically Meaningful? A Case Study in Medical Imaging
- 用梯度投影法分析ViT特征表示的语义性
- 微小扰动使分类准确率下降超60%,跨类别表征高度相似
- 警示医疗安全场景中使用ViT需谨慎,适合关注模型可信性的研究者
视觉变换器(ViTs)在疾病分类、分割和检测等医学影像任务中因精度优于传统深度学习模型而迅速兴起。然而,其庞大的规模和自注意力机制带来的复杂交互使其难以解释。尤其关键问题是:这些模型产生的表征是否具有语义意义?本文采用基于投影梯度的算法表明,ViT表征缺乏语义意义,对微小变化极为敏感。存在不可察觉的图像差异却导致表征截然不同;相反,本应属于不同语义类别的图像却有几乎相同的表征。这种脆弱性可能导致不可靠的分类结果,例如微小扰动使分类准确率下降超过60%。据我们所知,这是首个系统揭示医学图像分类中ViT表征缺乏语义意义的工作,揭示了其在安全关键系统中部署的重大挑战。
原文摘要 · Abstract (English)
Vision transformers (ViTs) have rapidly gained prominence in medical imaging tasks such as disease classification, segmentation, and detection due to their superior accuracy compared to conventional deep learning models. However, due to their size and complex interactions via the self-attention mechanism, they are not well understood. In particular, it is unclear whether the representations produced by such models are semantically meaningful. In this paper, using a projected gradient-based algorithm, we show that their representations are not semantically meaningful and they are inherently vulnerable to small changes. Images with imperceptible differences can have very different representations; on the other hand, images that should belong to different semantic classes can have nearly identical representations. Such vulnerability can lead to unreliable classification results; for example, unnoticeable changes cause the classification accuracy to be reduced by over 60\%. %. To the best of our knowledge, this is the first work to systematically demonstrate this fundamental lack of semantic meaningfulness in ViT representations for medical image classification, revealing a critical challenge for their deployment in safety-critical systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。