让大模型根据输入语义动态调整行为,无需训练即可提升表现
Semantics-Adaptive Activation Intervention for LLMs via Dynamic Steering Vectors
- 用对比输入的激活差异生成动态引导向量
- 在推理时按语义方向调节关键神经元,显著提升任务性能
- 适用于多种模型和任务,低成本且通用性强
大语言模型在众多任务中表现卓越,但对其行为进行对齐仍具挑战。激活干预作为一种高效经济的方法,可修改模型行为。然而,现有方法均使用固定引导向量,无法适应不同输入语义。为此,我们提出语义自适应动态干预(SADI),在推理时构建动态引导向量以干预模型激活。具体而言,SADI利用对比样本的激活差异,精准识别模型中关键组件(如注意力头、隐藏状态、神经元)用于靶向干预。推理时,SADI根据输入语义方向动态缩放逐元素激活,实现行为调节。实验表明,SADI显著优于现有基线,无需训练即可提升任务表现。其成本效益高,且在多种模型架构和任务上具备泛化能力,展现出作为通用对齐技术的巨大潜力。
原文摘要 · Abstract (English)
Large language models (LLMs) have achieved remarkable performance across many tasks, yet aligning them with desired behaviors remains challenging. Activation intervention has emerged as an effective and economical method to modify the behavior of LLMs. Despite considerable interest in this area, current intervention methods exclusively employ a fixed steering vector to modify model activations, lacking adaptability to diverse input semantics. To address this limitation, we propose Semantics-Adaptive Dynamic Intervention (SADI), a novel method that constructs a dynamic steering vector to intervene model activations at inference time. More specifically, SADI utilizes activation differences in contrastive pairs to precisely identify critical elements of an LLM (i.e., attention heads, hidden states, and neurons) for targeted intervention. During inference, SADI dynamically steers model behavior by scaling element-wise activations based on the directions of input semantics. Experimental results show that SADI outperforms established baselines by substantial margins, improving task performance without training. SADI's cost-effectiveness and generalizability across various LLM backbones and tasks highlight its potential as a versatile alignment technique.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。