通过几何对齐的稀疏编码电路,实现多层语义行为精准控制。
CircuitSteer: Geometrically Aligned Multi-Layer Steering via Sparse Autoencoder Circuits

- 基于稀疏自编码器识别跨层语义电路,构建几何对齐的特征流路径。
- 在毒性、情绪强度等任务上实现零退化流畅干预,优于现有方法。
- 适合需要高精度行为调控的研究者,尤其关注复杂行为如奉承与拒绝。
大型语言模型的行为控制仍是人工智能对齐的关键挑战。现有方法如对比激活添加(CAA)通常依赖从整体激活差异中提取的固定单层干预,对语义多样输入施加单一干预,且难以在多层间维持一致行为变化,限制了控制效果。本文提出CircuitSteer框架,利用稀疏自编码器(SAEs)识别分布于多层的连贯语义电路。通过构建基于特征共激活和解码方向几何对齐的特征流电路,分离出负责目标行为的特定多层子电路。随后从这些稀疏特征合成密集干预向量,并施加多点干预以引导模型内部语义轨迹。我们在涵盖毒性、情绪强度、奉承与拒绝等多样化任务的对比样本上评估该方法,覆盖两个模型家族。在所有模型和数据集上,CircuitSteer是唯一能持续生成保持流畅性的干预方案;其他方法或牺牲文本质量,或缺乏覆盖范围,在复杂行为如奉承与拒绝上完全失效。结果表明,通过强制选择特征间的几何对齐,实现的多层电路控制比静态单点干预更稳健有效。代码已开源:https://github.com/mehrshad-sdtn/CircuitSteer。
原文摘要 · Abstract (English)
Controlling the behavior of large language models (LLMs) remains a critical challenge for AI alignment. Existing steering methods, such as Contrastive Activation Addition (CAA), typically rely on fixed single-layer interventions derived from aggregate activation differences. These methods impose a single intervention across semantically diverse inputs and often fail to sustain consistent behavioral changes across layers, limiting the effectiveness of the steering. In this work, we introduce CircuitSteer, a novel framework that leverages Sparse Autoencoders (SAEs) to identify and manipulate coherent semantic circuits distributed across multiple layers. By constructing a feature flow circuit based on feature co-activation and the geometric alignment of decoder directions, we isolate the specific multi-layer subcircuits responsible for a target behavior. We then synthesize dense steering vectors from these sparse features and apply multi-point interventions to guide the model's internal semantic trajectory. We evaluate CircuitSteer using contrastive examples across a diverse set of tasks, including toxicity, emotion-intensity, sycophancy, and refusal, spanning two model families. Across all models and datasets, CircuitSteer is the only method to consistently produce fluency-preserving interventions; competing methods either sacrifice text quality or lack coverage, failing entirely on complex behaviors like sycophancy and refusal. These results demonstrate that multi-layer circuit steering, enabled by enforcing geometric alignment among selected features, yields strictly more robust and effective behavioral control than static single-point interventions. Code is available at https://github.com/mehrshad-sdtn/CircuitSteer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。