提出无梯度方法揭示视觉模型对图像变换的隐藏不变性。
Stretching Beyond the Obvious: A Gradient-Free Framework to Unveil the Hidden Landscape of Visual Invariance
- 通过双目标优化,系统寻找使神经元响应保持不变的图像变换。
- 发现深层特征不变性需更强纹理/姿态变化,且远超仿射变换范围。
- 揭示深度网络的不变性在人类可解释性上反而下降,适合研究模型泛化机制。
理解视觉单元编码的特征组合对于揭示图像如何转化为支持识别的表示至关重要。现有特征可视化方法通常只捕捉单元最敏感的图像,但无法揭示响应保持不变的变换流形,而这正是视觉泛化的核心。本文提出一种模型无关、无梯度的框架 Stretch-and-Squeeze (SnS),用于系统表征单元的最大不变刺激及其对抗扰动的脆弱性,适用于生物与人工视觉系统。SnS 将变换建模为双目标优化问题:为探测不变性,寻找能最大改变参考图像在特定处理阶段表示(拉伸),同时保持下游单元激活不变(压缩)的图像扰动;为探测对抗敏感性,则反向操作,最大化单元激活变化而最小化上游表示变化。应用于 CNN 时,发现其不变变换在像素空间中距离参考图像更远,且更强烈保持目标单元响应。不同层级的表示优化导致不同不变性:像素级变化主要影响亮度和对比度,中深层则主要改变纹理和姿态。通过测量分层不变图像在 L2 增强网络中被人类及其他观察者网络分类的准确率,发现深层拉伸后的图像在人类理解中可解释性显著下降,而标准模型则相反。
原文摘要 · Abstract (English)
Uncovering which feature combinations are encoded by visual units is critical to understanding how images are transformed into representations that support recognition. While existing feature visualization approaches typically infer a unit's most exciting images, this is insufficient to reveal the manifold of transformations under which responses remain invariant, which is critical to generalization in vision. Here we introduce Stretch-and-Squeeze (SnS), a model-agnostic, gradient-free framework to systematically characterize a unit's maximally invariant stimuli, and its vulnerability to adversarial perturbations, in both biological and artificial visual systems. SnS frames these transformations as bi-objective optimization problems. To probe invariance, SnS seeks image perturbations that maximally alter (stretch) the representation of a reference stimulus in a given processing stage while preserving unit activation downstream (squeeze). To probe adversarial sensitivity, stretching and squeezing are reversed to maximally perturb unit activation while minimizing changes to the upstream representation. Applied to CNNs, SnS revealed invariant transformations that were farther from a reference image in pixel-space than those produced by affine transformations, while more strongly preserving the target unit's response. The discovered invariant images differed depending on the stage of the image representation used for optimization: pixel-level changes primarily affected luminance and contrast, while stretching mid- and late-layer representations mainly altered texture and pose. By measuring how well the hierarchical invariant images obtained for L2 robust networks were classified by humans and other observer networks, we discovered a substantial drop in their interpretability when the representation was stretched in deep layers, while the opposite trend was found for standard models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。