arXiv:2607.24645cs.LGcs.AI2026-07

揭示稀疏自编码器特征如何影响模型输出,打破可解释性与可控性之间的误解

Sparse Autoencoders Encode Both Concepts and Functions: The Downstream Geometry of Feature Effects

论文配图:Sparse Autoencoders Encode Both Concepts and Functions: The Downstream Geometry of Feature Effects
图 1 · 摘自论文原文
  • 通过分析特征干预后模型输出变化的几何结构,发现多数特征无稳定作用方向
  • 区分出'值类'和'指针类'特征,前者更稳定,后者更分散
  • 为理解特征因果效应提供新视角,适合模型可解释性研究者

稀疏自编码器(SAEs)作为可解释性工具的广泛应用受限于其特征与模型行为之间不一致的关联。某些具有明确激活描述的特征可能产生微弱或意外的因果效应;操控效果在不同提示中变化不定,甚至与预期方向相反;基于激活的选择会遗漏能产生期望输出变化的特征。以往研究关注模型内部特征的几何结构,本文则聚焦特征干预引起的模型logits变化的几何特性。我们提出无监督框架FEGA,通过在多种上下文中移除同一活跃SAE特征,分析由此产生的logit变化云。结果显示,在不同SAE变体中,稳定的单维效应极为罕见:很少有特征具备可复用的方向性。为解释这种差异,我们区分了‘值类’特征(关联静态信息如事实属性)与‘指针类’特征(关联上下文依赖的操作)。值类特征更常表现出结构化、低维的效果,尽管通常跨越多个方向;而指针类特征则主要呈现弥散效应。结果表明,一个特征可以是可解释且具因果相关性,却无需提供稳定的操控方向。

原文摘要 · Abstract (English)

The wide-scale use of sparse autoencoders (SAEs) as interpretability tools is limited by inconsistent links between SAE features and model behavior. Features with clear activation descriptions may have weak or unexpected causal effects; steering can vary across prompts or oppose the intended direction; and activation-based feature selection can miss features that produce the desired output change. Prior work has studied feature geometry inside the model, where features are computed. We instead study the geometry of changes in model logits caused by feature interventions. We introduce Feature-Effect Geometry Analysis (FEGA), an unsupervised framework that removes the same active SAE feature across contexts and analyzes the resulting cloud of logit changes. Across SAE variants, consistent one-dimensional effects are rare: few features behave like reusable directions. To interpret this variation, we distinguish value-like features, tied to static information such as factual attributes, from pointer-like features, associated with context-dependent operations. Value-like features more often exhibit structured, low-dimensional effects, although these effects typically span several directions. Pointer-like features, by contrast, predominantly exhibit diffuse effects. Our results show that a feature can be interpretable and causally relevant without providing a stable direction for steering.

可解释性稀疏编码特征几何模型干预

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。