通过信号分析方法,发现注意力头可被精准调控以控制模型输出内容。
Head Pursuit: Probing Attention Specialization in Multimodal Transformers
- 用信号处理视角重解注意力探测,实现对注意力头的系统性排序。
- 仅修改1%精选注意力头,就能可靠抑制或增强特定概念输出。
- 适用于文本和图文任务,为大模型可解释性与可控性提供新工具。
语言和视觉-语言模型在众多任务中表现优异,但其内部机制仍不完全清楚。本文研究文本生成模型中各注意力头如何针对特定语义或视觉属性实现专业化。基于现有可解释性方法,我们将中间激活探测重新诠释为信号处理视角,使多样本分析更系统化,并按目标概念的相关性对注意力头进行排序。结果表明,单模态与多模态变换器在头层面均存在一致的专化模式。令人惊讶的是,仅通过我们方法选出的1%注意力头进行编辑,即可可靠地在模型输出中抑制或增强特定概念。该方法在问答、毒性缓解等语言任务,以及图像分类和图像描述等视觉-语言任务上得到验证。研究揭示了注意力层中可解释且可控制的结构,为理解与编辑大规模生成模型提供了简单有效工具。
原文摘要 · Abstract (English)
Language and vision-language models have shown impressive performance across a wide range of tasks, but their internal mechanisms remain only partly understood. In this work, we study how individual attention heads in text-generative models specialize in specific semantic or visual attributes. Building on an established interpretability method, we reinterpret the practice of probing intermediate activations with the final decoding layer through the lens of signal processing. This lets us analyze multiple samples in a principled way and rank attention heads based on their relevance to target concepts. Our results show consistent patterns of specialization at the head level across both unimodal and multimodal transformers. Remarkably, we find that editing as few as 1% of the heads, selected using our method, can reliably suppress or enhance targeted concepts in the model output. We validate our approach on language tasks such as question answering and toxicity mitigation, as well as vision-language tasks including image classification and captioning. Our findings highlight an interpretable and controllable structure within attention layers, offering simple tools for understanding and editing large-scale generative models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。