arXiv:2503.18762cs.LGcs.AI2025-03被引 5

解析视觉Transformer在扭曲图像中的注意力机制,揭示其对无关信息的处理方式。

Mechanistic Interpretability of Fine-Tuned Vision Transformers on Distorted Images: Decoding Attention Head Behavior for Transparent and Trustworthy AI

  • 通过机械可解释性分析注意力头行为,定位关键层与冗余模块。
  • 深层注意力头(如第6层)失效导致损失增加三倍,显示其任务重要性高。
  • 发现中间层注意力头专注信号区域,适合提升模型透明度与可信度。

机械可解释性提升了大型AI模型的安全性、可靠性和鲁棒性。本研究针对在包含非相关内容(如坐标轴标签、标题、颜色条)的扭曲二维光谱图上微调的视觉变换器(ViTs),分析了各注意力头对无关信息的处理机制。通过引入额外干扰特征,利用机械可解释性方法调试问题并揭示架构内在规律。注意力图评估各层头的贡献:早期层(1-3层)头影响极小,消融后均方误差(MSE)损失仅上升μ=0.11%(σ=0.09%),表明其关注低级非关键特征;而深层头(如第6层)损失增加三倍(μ=0.34%,σ=0.02%),显示更高任务重要性。中层(6-11层)表现出单义性,仅聚焦于啁啾信号区域;部分早期头(1-4层)虽为单义但无关任务(如文本检测、边缘/角点检测)。注意力图可区分单义头(精确定位啁啾区)与多义头(关注多个无关区域)。结果揭示了ViTs的功能专化特性,说明其如何处理相关与冗余信息。通过将变换器分解为可解释组件,本工作增强了模型理解,识别出潜在漏洞,推动更安全、透明的AI发展。

原文摘要 · Abstract (English)

Mechanistic interpretability improves the safety, reliability, and robustness of large AI models. This study examined individual attention heads in vision transformers (ViTs) fine tuned on distorted 2D spectrogram images containing non relevant content (axis labels, titles, color bars). By introducing extraneous features, the study analyzed how transformer components processed unrelated information, using mechanistic interpretability to debug issues and reveal insights into transformer architectures. Attention maps assessed head contributions across layers. Heads in early layers (1 to 3) showed minimal task impact with ablation increased MSE loss slightly (μ=0.11%, σ=0.09%), indicating focus on less critical low level features. In contrast, deeper heads (e.g., layer 6) caused a threefold higher loss increase (μ=0.34%, σ=0.02%), demonstrating greater task importance. Intermediate layers (6 to 11) exhibited monosemantic behavior, attending exclusively to chirp regions. Some early heads (1 to 4) were monosemantic but non task relevant (e.g. text detectors, edge or corner detectors). Attention maps distinguished monosemantic heads (precise chirp localization) from polysemantic heads (multiple irrelevant regions). These findings revealed functional specialization in ViTs, showing how heads processed relevant vs. extraneous information. By decomposing transformers into interpretable components, this work enhanced model understanding, identified vulnerabilities, and advanced safer, more transparent AI.

视觉Transformer可解释性注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。