arXiv:2602.02464cs.CL2026-02被引 7

用局部几何结构分解语言模型激活,捕捉非线性概念。

From Directions to Regions: Decomposing Activations in Language Models via Local Geometry

  • 用MFA模型将激活空间拆解为多个高斯区域及局部变化。
  • 在Llama-3.1-8B和Gemma-2-2B上验证了对复杂结构的捕捉能力。
  • 适合做概念发现与模型控制,优于传统方向方法。

语言模型中的激活分解方法通常依赖于激活空间中概念实现的几何假设。现有方法寻找单一全局方向,隐含线性可分性假设,忽略了具有非线性或高维结构的概念。本文采用可扩展的无监督方法——混合因子分析(MFA),将激活空间建模为多个具有局部协方差结构的高斯区域。MFA将激活分解为两部分:区域在激活空间中的中心点,以及相对于中心的局部变化。我们为Llama-3.1-8B和Gemma-2-2B训练了大规模MFA,并证明其能有效捕捉激活空间中的复杂非线性结构。在定位与引导任务上的评估显示,MFA优于无监督基线,性能接近有监督定位方法,且在多数情况下优于稀疏自编码器。结果表明,通过子空间表达的局部几何结构,是实现可扩展概念发现与模型控制的有力工具,能有效刻画孤立方向无法捕捉的复杂结构。

原文摘要 · Abstract (English)

Activation decomposition methods in language models are tightly coupled to geometric assumptions on how concepts are realized in activation space. Existing approaches search for individual global directions, implicitly assuming linear separability, which overlooks concepts with nonlinear or multi-dimensional structure. In this work, we leverage Mixture of Factor Analyzers (MFA) as a scalable, unsupervised alternative that models the activation space as a collection of Gaussian regions with their local covariance structure. MFA decomposes activations into two compositional geometric objects: the region's centroid in activation space, and the local variation from the centroid. We train large-scale MFAs for Llama-3.1-8B and Gemma-2-2B, and show they capture complex, nonlinear structures in activation space. Moreover, evaluations on localization and steering benchmarks show that MFA outperforms unsupervised baselines, is competitive with supervised localization methods, and often achieves stronger steering performance than sparse autoencoders. Together, our findings position local geometry, expressed through subspaces, as a promising unit of analysis for scalable concept discovery and model control, accounting for complex structures that isolated directions fail to capture.

激活分解局部几何语言模型MFA

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。