解析视觉模型决策过程,从传统CNN到大模型时代的解释方法演进
Class Activation Mapping in Explainable Computer Vision: A Method-Centered Review of CNN, Transformer, and Foundation-Model-Era Visual Explanations
- 按归因机制、架构依赖与评估目标分类,系统梳理57篇核心论文
- 揭示从单类低分辨率解释转向多层、概率化、注意力感知的复杂解释趋势
- 适合关注可解释AI、模型可信度及大模型视觉理解的研究者
类激活映射(CAM)是可解释人工智能中应用最广泛的可视化解释方法之一。其核心目标直观:将模型内部证据转化为热力图,突出支持目标类别或概念的图像区域、卷积通道、注意力标记或图像块。自2016年首个CAM提出以来,该领域已远超基于全局平均池化的CNN分类器。当前的CAM类方法涵盖梯度后处理、无梯度评分与消融法、高分辨率上采样、弱监督定位与分割、Transformer标记归因、因果与去偏方法,以及利用CLIP、DINO、SAM或特征分布对比的大模型时代方法。本文综述了2016年以来57篇以方法为中心的论文,构建了按归因机制、架构依赖和评估目标划分的分类体系,系统回顾了梯度型CAM、近期及混合型方法、以及模型/架构感知方法。研究发现,领域正从单一类别的低分辨率解释,转向比较性、多层、概率化、标记感知与大模型感知的复合解释。然而,评估标准仍碎片化,忠实度、定位精度、鲁棒性、计算成本与人类信任常采用不同协议。因此,本文不仅分析各方法贡献,更指出其遗留缺口及后续方法如何填补。
原文摘要 · Abstract (English)
Class activation mapping (CAM) is one of the most widely used visual explanation families in explainable artificial intelligence. Its purpose is intuitive: it converts internal model evidence into a heatmap that highlights the image regions, convolutional channels, tokens, or patches that support a target class or concept. Since the first CAM formulation in 2016, the field has moved far beyond global-average-pooled CNN classifiers. CAM-style methods now include gradient-based post-hoc explanations, gradient-free score and ablation methods, high-resolution upscaling, weakly supervised localization and segmentation, transformer token attribution, causal and debiasing methods, and foundation-model-era approaches that use CLIP, DINO, SAM, or feature-distribution comparisons. This review synthesizes a strict corpus of 57 method-centered papers published from 2016 onward. The paper develops a taxonomy that separates methods by attribution mechanism, architectural dependence, and evaluation objective. It then reviews gradient-based CAMs, recent and hybrid CAM-style methods, and model-based or architecture-aware methods. Across the corpus, the main trend is clear: the field is shifting from explaining one class score in one low-resolution CNN layer toward comparative, multi-layer, probabilistic, token-aware, and foundation-model-aware explanations. At the same time, evaluation remains fragmented. Faithfulness, localization, robustness, computational cost, and human trust are often measured with different protocols. The review therefore emphasizes not only what each method contributes, but also which gap it leaves open and which later methods attempt to close that gap.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。