通过引导梯度方向提升Transformer注意力解释能力,揭示模型决策机制。
Transformer Interpretability from Perspective of Attention and Gradient

- 用梯度方向指导注意力,实现更全面的特征区域解释
- 发现ViT与人类对图像感知差异,可微调分类且人眼难察觉
- 为理解模型内部机制提供新方法,适合关注AI可解释性研究者
尽管研究者更关注Transformer模型的性能表现,但其可解释性仍不可忽视。梯度被广泛用于Transformer的解释分析。本文从注意力与梯度的视角出发,深入研究了Transformer的解释问题,并提出一种通过引导梯度方向(即注意力方向)来实现解释的新方法。该方法能够更全面地识别特征区域,提供细节级解释,有助于深入理解Transformer的工作机制。利用视觉Transformer(ViT)与人类对图像感知方式的差异,我们实现了对图像类别的细微重写,该变化在人眼几乎无法察觉,可能在特定场景下带来安全风险。
原文摘要 · Abstract (English)
Although researchers' attention is more focused on the performance of Transformer models, the interpretation of Transformer can never be ignored. Gradient is widely utilized in Transformer interpretation. From the perspective of attention and gradient, we conduct an in-depth study of Transformer interpretation and propose a method to achieve it by guiding the gradient direction, or more precisely, the attention direction. The method enables more comprehensive interpretation of feature regions, offers detail interpretation, and helps to better understand Transformer mechanism. Leveraging the difference in how Vision Transformer (ViT) and humans perceive images, we alter the class of an image in a way that is almost imperceptible to the human eye. This class rewriting phenomenon may potentially pose security risks in certain scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。