arXiv:2605.31183cs.CLcs.AI2026-05

改进的稀疏自编码器可有效控制大模型输出,表现接近LoRA。

Steering LLMs? Actually, Sparse Autoencoders can outperform simple baselines

论文配图:Steering LLMs? Actually, Sparse Autoencoders can outperform simple baselines
图 1 · 摘自论文原文
  • 用监督式管道筛选并标注特征,提升SAE控制能力。
  • 在AxBench上性能接近参考LoRA,验证了有效性。
  • 高稀疏性非必要,适合想解释模型行为的研究者。

稀疏自编码器(SAEs)被视为探索大语言模型内部机制和控制生成输出的有前景方向。然而,在Wu等(2025)提出的AxBench基准测试中,SAEs的表现未达预期,低于若干简单基线。本文对此提出部分反驳,表明当使用我们提出的监督式特征选择与标注流程时,SAEs在AxBench上的表现可接近甚至媲美参考的LoRA方法。此外,仅依赖可解释性组件,所选特征与其标签之间表现出惊人因果关系。最后,我们发现高稀疏性(低l0)并非成功可解释性控制的关键,这与Wang等(2025)的早期结论相悖。

原文摘要 · Abstract (English)

Sparse Autoencoders (SAEs) have been seen as a promising avenue for exploring the internals of Large Language Models (LLMs) and for steering model output generation. When AxBench - a model steering benchmark - was introduced in Wu et al. (2025), SAEs did not seem to live up to their original hype due to poor steering performance relative to a set of simple baselines. This work serves as a partial rebuttal for Sparse Autoencoders and suggests that the results of Wu et al. (2025) did not do them full justice. We find that Sparse Autoencoders can, in fact, perform close to on par with the reference LoRA performance on the AxBench benchmark, when features are selected and labelled with our supervised pipeline. We also find that our pipeline selects features that are surprisingly causal of their identified labels when using only its interpretability-based components. Lastly, we present evidence that high sparsity (low l0) may not be crucial for successful steering based on interpretability, which is in contrast to the earlier findings in Wang et al. (2025).

稀疏自编码器模型控制可解释性大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。