arXiv:2608.10537cs.AIcs.LG2026-08

用非局部性量化SAE特征语义抽象程度,区分上下文推理与词元驱动特征。

Measuring Semantic Abstractness of SAE Features via Nonlocality

论文配图:Measuring Semantic Abstractness of SAE Features via Nonlocality
图 1 · 摘自论文原文
  • 提出特征非局部性(FNL),通过位置影响熵衡量特征抽象度。
  • 73%~84%准确区分上下文特征与词元特征,高FNL对应更高层次语义。
  • 可无标签评估模型解释有效性,适合机制分析与干预特征筛选。

稀疏自编码器(SAEs)已帮助揭示大语言模型行为的机制解释,如推理和越狱等,通过理解相关任务敏感且因果有效的特征。然而,现有基于自动解释的语义描述或因果控制能力无法完全判断特征的抽象层次。为此,我们引入特征非局部性(FNL),定义为标准化后各位置对SAE特征激活的影响熵。实验显示,FNL与现有基于LLM的特征抽象度代理指标高度相关,并能有效区分依赖上下文的推理特征与仅依赖词元的特征,在随机抽取的一对上下文特征与词元特征中,正确识别出上下文特征具有更高FNL的比例达73%~84%。我们展示了两个下游应用:审计用于越狱缓解的SAE特征,发现最有效的特征多为位置特征且FNL低,而非真正识别有害意图;在DeepSeek-R1-Distill-Llama-8B上,控制高FNL特征使MATH-500得分提升4.6分,优于低FNL特征,但增益具模型依赖性。结论:FNL提供一种不依赖LLM、无需标签、基于相关性的特征抽象度见证,适用于评估机制解释及选择下游干预特征。

原文摘要 · Abstract (English)

Sparse autoencoders (SAEs) have helped uncover mechanistic explanations for LLM behaviours such as reasoning, jailbreaking etc., via understanding the corresponding task-relevant and causally effective features. To evaluate such mechanistic explanations, downstream studies must distinguish surface lexical features from genuinely high-level ones. However, neither an autointerp-based semantic description nor causal steering utility fully resolves the abstraction level of a feature. To this end, we introduce \emph{Feature Nonlocality} (FNL), defined as the entropy of the normalized per-position influence on an SAE feature's activation. We report that FNL correlates with existing LLM-based proxy metrics of feature semantic abstractness, and successfully distinguishes context-dependent reasoning features from token-driven ones, correctly assigning the higher FNL to the contextual feature in $73$--$84\%$ of randomly drawn pairs that consist of one contextual and one token-level feature. We demonstrate two downstream applications. We audit SAE-based features used for jailbreak mitigation and find surprisingly that most effective features are positional features with low FNL rather than genuinely recognizing harmful intents. We report that steering high-FNL features in DeepSeek-R1-Distill-Llama-8B improves MATH-500 accuracy by $4.6$ points over the unsteered model and outperforms steering low-FNL features, though the gains are model-specific. We conclude that FNL provides an LLM-independent, label-free, correlational witness of the abstraction level of an SAE feature, with applications in evaluating mechanistic explanations as well as selecting features for downstream interventions.

SAE抽象度机制解释特征审计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。