让激活调控模仿提示引导,提升模型推理时的可控性。
Steer Like the LLM: Activation Steering that Mimics Prompting

- 从激活中自动生成针对性调控系数,模拟提示的效果
- 在三个基准上超越现有激活调控方法,尤其保持高连贯性
- 适合需要精确控制生成内容的研究者和开发者
大型语言模型可通过提示或激活干预在推理阶段进行调控,但现有激活调控方法表现通常不如提示法。本文将提示调控建模为一种激活调控形式,探索是否可通过训练简单可解释模型来复现成功提示调控行为。分析发现,主流激活调控方法未能忠实模拟提示调控机制——后者对部分词元施加强干预,而几乎不影响其他词元。基于此,提出提示调控替代模型(PSR),从激活中估计词元级调控系数,并训练其模仿提示干预效果。在多个语言模型上的三项调控基准测试表明,PSR在保持高连贯性生成的前提下显著优于现有激活调控方法,在AxBench和人物设定调控任务中也媲美提示法。
原文摘要 · Abstract (English)
Large language models can be steered at inference time through prompting or activation interventions, but activation steering methods often underperform compared to prompt-based approaches. We propose a framework that formulates prompt steering as a form of activation steering and investigates whether distilling successful prompt steering behavior into simpler, interpretable models can close this gap. Our analysis reveals that popular activation steering methods are not faithful to the mechanics of prompt steering, which applies strong interventions on some tokens while barely affecting others. Based on these insights, we introduce Prompt Steering Replacement (PSR) models that estimate token-specific steering coefficients from the activations themselves and are trained to imitate prompt-based interventions. Experiments on three steering benchmarks across multiple language models show that PSR models outperform existing activation steering methods, especially when controlling for high-coherence completions, and also compare favorably to prompting on AxBench and persona steering.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。