用扩散模型将特征图精准还原为图像,助力理解神经网络内部机制
FeatInv: Spatially resolved mapping from feature space to input space using conditional diffusion models
- 基于条件扩散模型,实现从特征空间到输入空间的像素级映射
- 在多种图像分类模型上实现高质量图像重建,保持细节与结构
- 适合研究模型可解释性、概念操控与特征空间组成的研究者
内部表示对理解深度神经网络的特性与推理模式至关重要,但其可解释性仍具挑战。现有方法常依赖粗略近似来实现从特征空间到输入空间的映射。本文提出一种基于条件扩散模型的方法——使用预训练的高保真扩散模型,以空间分辨的特征图为条件,概率化地学习该映射。我们在多种预训练图像分类模型(从CNN到ViT)上验证了该方法的可行性,展示了出色的重建能力。通过定性对比与鲁棒性分析,证实了方法有效性,并展示了其在输入空间中概念操控可视化及特征空间复合性质探究等应用潜力。该方法在提升计算机视觉模型特征空间理解方面具有广泛前景。
原文摘要 · Abstract (English)
Internal representations are crucial for understanding deep neural networks, such as their properties and reasoning patterns, but remain difficult to interpret. While mapping from feature space to input space aids in interpreting the former, existing approaches often rely on crude approximations. We propose using a conditional diffusion model - a pretrained high-fidelity diffusion model conditioned on spatially resolved feature maps - to learn such a mapping in a probabilistic manner. We demonstrate the feasibility of this approach across various pretrained image classifiers from CNNs to ViTs, showing excellent reconstruction capabilities. Through qualitative comparisons and robustness analysis, we validate our method and showcase possible applications, such as the visualization of concept steering in input space or investigations of the composite nature of the feature space. This approach has broad potential for improving feature space understanding in computer vision models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。