动态调整提示词的激活强度,让大模型更精准地响应指令。
Contextual Linear Activation Steering of Language Models
- 根据上下文自动调节激活控制强度,避免统一力度导致效果不一。
- 在11个评测任务中优于传统方法,限数据下媲美甚至超越ReFT和LoRA。
- 适合需要精准行为定制且标注数据少的场景,可解释性强。
线性激活控制是一种有效激发大语言模型能力并用少量标注数据定制其行为的方法。然而,现有方法对所有令牌使用固定控制强度,导致不同输入提示下的控制效果不一致。本文提出上下文感知的线性激活控制(CLAS),动态适应不同上下文的控制强度。在11个控制基准和4种模型家族上,CLAS持续优于标准线性激活控制,并在有限标注数据设置下达到或超过ReFT和LoRA的性能。因此,我们建议将CLAS作为一种可扩展、可解释且精确的大语言模型定制与控制方法。
原文摘要 · Abstract (English)
Linear activation steering is a powerful approach for eliciting the capabilities of large language models and specializing their behavior using limited labeled data. While effective, existing methods often apply a fixed steering strength to all tokens, resulting in inconsistent steering quality across diverse input prompts. In this work, we introduce Contextual Linear Activation Steering (CLAS), a method that dynamically adapts linear activation steering to context-dependent steering strengths. Across eleven steering benchmarks and four model families, it consistently outperforms standard linear activation steering and matches or exceeds the performance of ReFT and LoRA in settings with limited labeled data. We therefore propose CLAS as a scalable, interpretable, and accurate method for specializing and steering large language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。