Transformer通过聚焦高激活词元,提升概念识别的可靠性。
The SuperActivator Mechanism: Transformers Concentrate Reliable Concept Signals in the Tail
- 发现注意力头会放大概念激活差异,使关键词元更突出。
- 仅用5%-10%的高激活词元即可实现最优检测效果,F1提升最高0.14。
- 适合需要可解释性与精准定位的模型分析场景。
概念向量旨在通过连接内部表示与人类可理解语义来提升模型可解释性,但其实际效用常受限于噪声和不一致的激活。本文揭示了超激活机制:Transformer动态地放大概念激活差距,将最可靠的概念证据集中在少数高激活词元中。理论证明,对齐概念的注意力头会乘法放大成对激活差距,已有极端激活增长最快。实证发现,尽管概念内与外激活分布重叠显著,但概念内分布形成明显正尾部,与噪声分离。这些高尾部词元被称为超激活器(SuperActivators),在概念正样本中具有一致性,成为概念存在的可靠指标。基于超激活器的检测方法在图像与文本模态、不同模型、层及提取技术下,相比标准激活聚合与提示基线,F1最高提升0.14,验证了洞察的普适性与实用性。进一步分析表明,最可靠的超激活器稀疏,通常仅需5%-10%的概念内词元激活即可达到最佳检测性能,且比全局概念向量捕捉更忠实的局部语义。
原文摘要 · Abstract (English)
Concept vectors aim to enhance model interpretability by linking internal representations with human-understandable semantics, but their practical utility is often limited by noisy and inconsistent activations. In this work, we uncover the SuperActivator Mechanism: a transformer dynamic that amplifies concept activation gaps, concentrating the most reliable concept evidence into a small set of high-activation tokens. To develop a theoretical understanding of this mechanism, we prove that concept-aligned attention heads multiplicatively amplify pairwise activation gaps, with already-extreme activations growing fastest. We find that this amplification is not just theoretical, but also occurs empirically on large-scale models: while in- and out-of-concept activation distributions overlap considerably, the in-concept distribution develops a positive tail clearly separated from the noise. These high-tail tokens, which we call SuperActivators, appear consistently across concept-positive samples, making them reliable indicators of concept presence. Accordingly, SuperActivator-based detection improves F1 by up to 0.14 over standard concept activation aggregators and prompting baselines across image and text modalities, models, layers, and concept extraction techniques, demonstrating the generality and practicality of our insights. Further empirical analysis demonstrates that the most reliable SuperActivators are sparse, with detection typically peaking when using only 5-10% of in-concept token activations, and capture more faithful localized semantics than global concept vectors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。