arXiv:2510.01246cs.CL2025-10被引 2

用稀疏自编码器单维度控制模型推理,提升数学题解题质量

A Comparative Analysis of Sparse Autoencoder and Activation Difference in Language Model Steering

  • 选最相关的单一编码维度进行控制,避免无关标点干扰
  • 引入逐词衰减策略,解决重复输出问题,使结果更稳定
  • 在数学推理任务上优于均值激活差方法,适合可控生成研究

稀疏自编码器(SAEs)已成为语言模型调控的有力工具。以往研究多采用前k个最具代表性的编码维度进行调控,但我们发现其中许多维度捕捉的是标点等非语义特征,而非指令等语义属性。为此,我们聚焦于最具相关性的单一编码维度(top-1),剔除冗余信息。进一步发现,固定值调控常导致输出退化,如重复单一词汇。为此,我们提出一种逐令牌衰减的调控策略,使与均值激活差异基线的对比更为可信。实验证明,对与推理相关的SAE维度进行调控,能可靠激发逐步数学推理过程,并提升推理质量,功能上类似添加引导标记。结果表明,在数学推理基准测试中,SAEs表现优于均值激活差异方法,在IF-Eval上性能相当。

原文摘要 · Abstract (English)

Sparse autoencoders (SAEs) have recently emerged as a powerful tool for language model steering. Prior work has explored top-k SAE latents for steering, but we observe that many dimensions among the top-k latents capture non-semantic features such as punctuation rather than semantic attributes like instructions. To address this, we propose focusing on a single, most relevant SAE latent (top-1), eliminating redundant features. We further identify a limitation in constant SAE steering, which often produces degenerate outputs such as repetitive single words. To mitigate this, we introduce a token-wise decaying steering strategy, enabling more faithful comparisons with mean activation difference baselines. Empirically, we show that steering an SAE latent associated with reasoning reliably elicits step-by-step mathematical reasoning and enhances inference quality, functionally resembling the effect of appending a guiding token. Our results demonstrate that SAEs outperform mean activation difference methods on mathematical reasoning benchmarks and match their performance on IF-Eval.

语言模型稀疏编码推理增强可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。