arXiv:2607.05615cs.LG2026-07

只对一半令牌施加控制,也能有效调节大模型行为。

A Coin Flip Per Token: Bernoulli Sparse Steering of Large Language Models

  • 随机选择部分令牌进行稀疏干预,降低对流畅性的影响
  • 仅干预50%令牌即可恢复密集干预的大部分效果
  • 适合希望低代价控制模型行为的研究者

通过稀疏自动编码器(SAEs)进行激活调控,可在不微调任务的情况下控制大语言模型行为,但传统方法在每个生成令牌上施加调控信号,持续扰动可能损害流畅性。本文提出两种新方法:随机令牌调控(STS)按概率p独立开启每令牌调控,以及随机块调控(SBS)每序列仅对前窗一次调控;两者均无需奖励模型或学习门控策略。在两个模型家族和两个行为任务中,仅干预50%的令牌即可恢复绝大部分密集调控效果,而30%的干预率已超越提示工程控制。最优调控强度与干预比例成反比,表明SAE调控受速率限制:行为结果取决于序列中累积的信号剂量。

原文摘要 · Abstract (English)

Activation steering via sparse autoencoders (SAEs) enables behavioral control of large language models without task-specific fine-tuning, but standard methods apply the steering signal at every generated token, incurring constant per-token perturbation that risks degrading fluency. We ask: is dense intervention necessary? We introduce Stochastic Token Steering (STS), which gates each token independently with probability $p$, and Stochastic Block Steering (SBS), which gates a leading window once per sequence; neither requires a reward model or learned gating policy. Across two model families and two behavioral tasks, steering only 50% of the tokens recovers most of the dense-steering effect while preserving fluency, and steering as few as 30% surpasses prompt-based control. The optimal steering magnitude scales inversely with the intervention ratio, revealing that SAE-mediated control is rate-limited: the behavioral outcome depends on cumulative signal dosage across a sequence.

模型控制稀疏调控语言模型生成质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。