通过激活引导发现的神经元能清晰揭示大模型对概念的理解。
ExpertLens: Activation steering features are highly interpretable
- 用激活引导方法定位特定概念的神经元,实现可解释性分析。
- 专家视角(ExpertLens)与人类认知高度一致,超越词向量对齐效果。
- 适用于想理解模型内部表征的研究者,轻量高效易部署。
大型语言模型中的激活引导方法已成为无需大量微调数据即可优化生成语言的有效手段。本文探究了激活引导所发现特征的可解释性:利用激活引导研究中的“寻找专家”方法,识别出负责特定概念(如‘cat’)的神经元,并通过专家视角(ExpertLens)检验这些神经元的表征。结果表明,ExpertLens在不同模型和数据集上具有稳定性,其表征与基于行为数据推断的人类认知高度吻合,达到人与人之间的对齐水平。相比词/句嵌入,ExpertLens显著提升对齐程度。通过重构人类概念组织结构,验证了其能提供大模型概念表征的精细视图。研究显示,ExpertLens是一种灵活且轻量的模型表征捕捉与分析方法。
原文摘要 · Abstract (English)
Activation steering methods in large language models (LLMs) have emerged as an effective way to perform targeted updates to enhance generated language without requiring large amounts of adaptation data. We ask whether the features discovered by activation steering methods are interpretable. We identify neurons responsible for specific concepts (e.g., ``cat'') using the ``finding experts'' method from research on activation steering and show that the ExpertLens, i.e., inspection of these neurons provides insights about model representation. We find that ExpertLens representations are stable across models and datasets and closely align with human representations inferred from behavioral data, matching inter-human alignment levels. ExpertLens significantly outperforms the alignment captured by word/sentence embeddings. By reconstructing human concept organization through ExpertLens, we show that it enables a granular view of LLM concept representation. Our findings suggest that ExpertLens is a flexible and lightweight approach for capturing and analyzing model representations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。