arXiv:2604.14477cs.AI2026-04被引 1

发现视觉Transformer内部的可解释计算通路,让模型决策过程更透明。

Seeing Through Circuits: Faithful Mechanistic Interpretability for Vision Transformers

论文配图:Seeing Through Circuits: Faithful Mechanistic Interpretability for Vision Transformers
图 1 · 摘自论文原文
  • 基于边的电路分析法,定位视觉Transformer中关键信息路径。
  • 成功识别出分类任务和对抗攻击中的核心计算通路。
  • 适合关注模型可解释性与安全性的研究人员使用。

神经网络内部推理过程的透明性是可解释性研究的核心,有助于提升模型的信任度、安全性与理解力。近年来,机制可解释性聚焦于任务特定的计算图,即模型组件间的连接(边)构成的电路。此类基于边的电路已在大语言模型中被定义,但现有视觉方法仅关注基于神经元的电路,仅能说明信息编码,无法揭示其在复杂网络中的传递路径。本文首次探索在视觉Transformer中通过计算图识别有效机制电路。提出自动视觉电路发现(Vi-CD)方法,成功恢复分类任务中的类别特异性电路,识别出CLIP模型中导致字体攻击的电路,并发现可引导纠正有害行为的电路。结果表明,从视觉Transformer中可恢复出具有洞察力且可操作的边级电路,显著增强模型内部计算的透明性。

原文摘要 · Abstract (English)

Transparency of neural networks' internal reasoning is at the heart of interpretability research, adding to trust, safety, and understanding of these models. The field of mechanistic interpretability has recently focused on studying task-specific computational graphs, defined by connections (edges) between model components. Such edge-based circuits have been defined in the context of large language models, yet vision-based approaches so far only consider neuron-based circuits. These tell which information is encoded, but not how it is routed through the complex wiring of a neural network. In this work, we investigate whether useful mechanistic circuits can be identified through computational graphs in vision transformers. We propose an effective method for Automatic Visual Circuit Discovery (Vi-CD) that recovers class-specific circuits for classification, identifies circuits underlying typographic attacks in CLIP, and discovers circuits that lend themselves for steering to correct harmful model behavior. Overall, we find that insightful and actionable edge-based circuits can be recovered from vision transformers, adding transparency to the internal computations of these models.

可解释性视觉Transformer电路分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。