arXiv:2508.07281cs.CVcs.AI2025-08被引 1

用激活最大化方法可视化深度模型内部特征,揭示其层次结构与潜在漏洞。

Representation Understanding via Activation Maximization

  • 提出统一框架,适用于CNN与ViT的中间层特征可视化。
  • 在传统CNN和现代ViT中均验证了方法的有效性。
  • 发现可利用该方法生成对抗样本,揭示模型决策边界。

理解深度神经网络(DNN)的内部特征表示是实现模型可解释性的基础。受神经科学中通过视觉刺激探测生物神经元的启发,近期研究采用激活最大化(AM)方法合成能激发人工神经元强响应的输入。本文提出一个统一的特征可视化框架,适用于卷积神经网络(CNN)和视觉变换器(ViT)。不同于以往主要关注CNN最后一层输出神经元的研究,我们将特征可视化扩展至中间层,深入揭示了学习特征表示的层级结构。此外,我们研究了激活最大化如何用于生成对抗样本,揭示了DNN的潜在漏洞和决策边界。实验表明,该方法在传统CNN和现代ViT中均有效,展现出良好的泛化能力和解释价值。

原文摘要 · Abstract (English)

Understanding internal feature representations of deep neural networks (DNNs) is a fundamental step toward model interpretability. Inspired by neuroscience methods that probe biological neurons using visual stimuli, recent deep learning studies have employed Activation Maximization (AM) to synthesize inputs that elicit strong responses from artificial neurons. In this work, we propose a unified feature visualization framework applicable to both Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs). Unlike prior efforts that predominantly focus on the last output-layer neurons in CNNs, we extend feature visualization to intermediate layers as well, offering deeper insights into the hierarchical structure of learned feature representations. Furthermore, we investigate how activation maximization can be leveraged to generate adversarial examples, revealing potential vulnerabilities and decision boundaries of DNNs. Our experiments demonstrate the effectiveness of our approach in both traditional CNNs and modern ViT, highlighting its generalizability and interpretive value.

特征可视化可解释性ViT对抗攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。