用最大似然估计推导大模型概念方向,更灵活高效。
Toward a Flexible Framework for Linear Representation Hypothesis Using Maximum Likelihood Estimation
- 将概念视为单位向量,通过激活差建模为vMF分布求方向。
- 在多个任务中优于现有方法,尤其在监控与操控上表现突出。
- 不依赖单标记反事实对,适合复杂语义概念分析。
线性表征假说认为高级概念在大模型表征空间中以线性方向编码。Park等(2024)通过因果内积统一了多种解释,但其框架依赖单标记反事实对,难以处理模糊对比对,限制了复杂或上下文相关概念的应用。本文提出将二元概念视为标准表征空间中的单位向量,利用大模型的(神经)激活差异与最大似然估计(MLE)计算概念方向(即控制向量)。所提方法SAND将激活差建模为冯·米塞斯-费舍尔(vMF)分布样本,提供了一种严谨的推导方式。我们扩展了Park等(2024)的适用性,消除了对解嵌入表示和单标记对的依赖。在LLaMA系列模型上针对多样概念与基准的实验表明,该轻量级方法具有更高灵活性,在激活工程任务(如监控与操控)中表现更优。
原文摘要 · Abstract (English)
Linear representation hypothesis posits that high-level concepts are encoded as linear directions in the representation spaces of LLMs. Park et al. (2024) formalize this notion by unifying multiple interpretations of linear representation, such as 1-dimensional subspace representation and interventions, using a causal inner product. However, their framework relies on single-token counterfactual pairs and cannot handle ambiguous contrasting pairs, limiting its applicability to complex or context-dependent concepts. We introduce a new notion of binary concepts as unit vectors in a canonical representation space, and utilize LLMs' (neural) activation differences along with maximum likelihood estimation (MLE) to compute concept directions (i.e., steering vectors). Our method, Sum of Activation-base Normalized Difference (SAND), formalizes the use of activation differences modeled as samples from a von Mises-Fisher (vMF) distribution, providing a principled approach to derive concept directions. We extend the applicability of Park et al. (2024) by eliminating the dependency on unembedding representations and single-token pairs. Through experiments with LLaMA models across diverse concepts and benchmarks, we demonstrate that our lightweight approach offers greater flexibility, superior performance in activation engineering tasks like monitoring and manipulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。