arXiv:2601.22447cs.LG2026-01

用权重直接分析稀疏自编码器特征,揭示其真实功能

Beyond Activation Patterns: A Weight-Based Out-of-Context Explanation of Sparse Autoencoder Features

  • 基于模型权重而非激活值,直接测量特征的函数作用
  • 四分之一特征能直接预测输出词,深度依赖注意力结构
  • 语义与非语义特征在注意力电路中分布差异明显

稀疏自编码器(SAEs)已成为解析语言模型表征的有力工具。现有方法通过激活模式推断特征语义,却忽略了特征被训练用于重建前向传播中的计算性激活。本文提出一种全新的基于权重的解释框架,通过直接分析权重交互来衡量特征的功能影响,无需依赖激活数据。在Gemma-2和Llama-3.1模型上的三项实验表明:(1)1/4的特征能直接预测输出词;(2)特征以深度依赖的方式参与注意力机制;(3)语义与非语义特征在注意力回路中具有不同的分布特征。该分析填补了SAE特征可解释性中“无上下文”部分的缺失。

原文摘要 · Abstract (English)

Sparse autoencoders (SAEs) have emerged as a powerful technique for decomposing language model representations into interpretable features. Current interpretation methods infer feature semantics from activation patterns, but overlook that features are trained to reconstruct activations that serve computational roles in the forward pass. We introduce a novel weight-based interpretation framework that measures functional effects through direct weight interactions, requiring no activation data. Through three experiments on Gemma-2 and Llama-3.1 models, we demonstrate that (1) 1/4 of features directly predict output tokens, (2) features actively participate in attention mechanisms with depth-dependent structure, and (3) semantic and non-semantic feature populations exhibit distinct distribution profiles in attention circuits. Our analysis provides the missing out-of-context half of SAE feature interpretability.

稀疏自编码器模型解释注意力机制权重分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。