通过注意力层面干预提升语言模型行为控制的稳定性与连贯性
Prompt-Activation Duality: Improving Activation Steering via Attention-Level Interventions

- 从系统提示词中提取控制信号,按令牌级门控注入注意力机制
- 多轮对话中平均连贯性漂移从-18.6提升至-1.9,第10轮角色表达达93.1
- 适用于需长期保持角色一致性的对话系统,如虚拟助手、角色扮演
激活控制在推理时通过向内部表示添加方向来调节语言模型行为,但标准残差流控制在状态化对话中可能失效。我们发现键值缓存污染是关键失败原因:被操控的令牌状态被存储并重复使用,导致局部扰动演变为累积性连贯性退化。为此,我们提出门控裁剪注意力差分控制(GCAD),从系统提示词对自注意力的贡献中提取控制信号,并以令牌级门控方式施加。在角色控制实验中,GCAD在保持性格控制的同时显著提升长程连贯性。在主多轮基准测试中,平均连贯性漂移从-18.6改善至-1.9,第10轮角色表达从78.0提升至93.1。结果表明,当干预遵循模型原本用于行为控制的提示引导路径时,激活控制更具可靠性。
原文摘要 · Abstract (English)
Activation steering controls language model behavior by adding directions to internal representations at inference time, but standard residual-stream steering can fail in stateful dialogue. We identify KV-cache contamination as a key failure mode: steered token states are stored and repeatedly reused, turning a local perturbation into cumulative coherence degradation. To address this challenge, we propose Gated Cropped Attention-Delta steering (GCAD), which extracts steering signals from system-prompt contributions to self-attention and applies them with token-level gating. Across persona-steering experiments, GCAD preserves trait control while substantially improving long-horizon coherence. On the main multi-turn benchmark, GCAD improves average coherence drift from -18.6 to -1.9 and raises turn-10 trait expression from 78.0 to 93.1. These results suggest that activation steering becomes more reliable when interventions follow the prompt-mediated pathways that models already use for behavioral control.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。