通过单向编辑实现大模型风格精准控制,无需训练
Controlling Chat Style in Language Models via Single-Direction Editing
- 将风格视为激活空间中的线性方向,实现可解释的风格编辑
- 在十余个模型上验证,风格遵循度高且保持核心能力
- 轻量无训练,支持风格组合与不良行为消除,适合实用部署
大型语言模型(LLM)中的风格控制仍具挑战性,现有方法多依赖提示工程或后训练对齐。本文从表征工程视角出发,验证了情感基调、语言结构等不同风格属性在模型激活空间中可能以线性方向编码的假设。我们在多种风格上提供了强有力的实证证据,并基于此提出一种无需训练、轻量高效的风格控制方法。该方法支持线性风格组合,可通过消融抑制不良行为;实验在十余个模型上验证,均实现了高风格遵循度,同时保留核心语言能力,计算开销极低。
原文摘要 · Abstract (English)
Controlling stylistic attributes in large language models (LLMs) remains challenging, with existing approaches relying on either prompt engineering or post-training alignment. This paper investigates this challenge through the lens of representation engineering, testing the hypothesis that distinct stylistic attributes - from emotional tone to linguistic structure - are encoded as linear directions in the model's activation space. We provide strong empirical evidence for this hypothesis across a wide range of styles and, based on this finding, present a lightweight, training-free method for precise style control. Our approach supports linear style composition, enhances safety by ablating undesirable behaviors, and, as confirmed by experiments on over a dozen models, achieves high style adherence while preserving core capabilities at minimal computational cost.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。