用稀疏编码让大模型行为更可控,精准调节特定功能。
Steering Large Language Model Activations in Sparse Spaces
- 通过稀疏自编码器提取独立语义特征,实现行为精准控制。
- 在Gemma 2上验证,可实现细粒度行为调制,效果随模型规模提升。
- 适合需要可解释干预的对齐研究,尤其关注行为可控性场景。
AI对齐的关键挑战在于如何在推理阶段引导大语言模型(LLM)遵循期望行为。激活量调节通过修改推理时的内部激活提供了一种潜在解决方案。然而,以往在密集激活空间中的方法受限于超位置现象,即多个特征相互纠缠,导致可解释性差、控制不精确。相比之下,稀疏表示为更可解释的行为调控提供了新机会。本文提出稀疏激活调节(SAS),利用稀疏自编码器(SAEs)在稀疏空间中调控LLM行为。通过对比提示配对策略,分离出与特定行为相关的特征,从而实现行为的定向增强或抑制。在Gemma 2 LLM上的实验表明,SAS向量支持细腻的行为调节与更精细的控制。此外,扩大SAE规模能提升SAS向量的单义性,意味着干预更可靠、更可解释。
原文摘要 · Abstract (English)
A key challenge in AI alignment is guiding large language models (LLMs) to follow desired behaviors at test time. Activation steering, which modifies internal model activations during inference, offers a potential solution. However, prior work in dense activation spaces struggles with superposition, wherein multiple features become entangled, limiting interpretability and precise control. In contrast, sparse representations provide an untapped opportunity for more interpretable behavior modulation. In this work, we introduce sparse activation steering (SAS), a method that leverages sparse autoencoders (SAEs) to steer LLM behavior in sparse spaces. By isolating behavior-specific features through a contrastive prompt-pairing approach, we define a set of features that can selectively reinforce or suppress behaviors. Experiments on Gemma 2 LLMs show that SAS vectors enable nuanced behavioral modulation and finer-grained control. Furthermore, scaling SAEs improves monosemanticity of SAS vectors, suggesting more reliable and interpretable interventions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。