用激活空间的对比引导,实现多维度作者风格精准控制。
Can Activation Steering Capture Multidimensional Authorship Style?

- 通过修辞维度的对比提示,在激活空间构建风格表示。
- 不同风格方向共享核心特征,但特定残差携带真实风格信号。
- 无需训练,适合跨领域风格迁移与可控生成任务。
激活转向在控制大语言模型生成方面展现出潜力,但其能否处理多维且难以定义的作者风格仍不明确。本文探讨了基于修辞动机的结构化对比提示是否能在激活空间中直接构建丰富的风格表征,而无需自然语言描述或专门训练。结果表明,这些方向共享一个共同的作者风格主干,但在特定维度的残差上存在冲突,这些残差承载着真实的风格信息,解释了为何简单聚合会失败。为此,我们提出了无训练的逐方面激活转向(A3S)框架,该框架融合各方面的对比方向,并采用干扰感知聚合,按实例调整转向强度。A3S在真正多方面的作者风格迁移中表现更优,在跨领域基准上的偏好评估中超越有训练基线,且始终保持目标样本与示例之间的重叠度低。
原文摘要 · Abstract (English)
Activation steering has shown promise for controlling LLM generation along well-defined attributes, but it remains unclear whether it can handle the multidimensional and hard-to-define nature of authorship style. We ask whether structured contrastive prompting along rhetorically-motivated dimensions can construct rich style representations directly in activation space, bypassing the need for natural language style descriptors or dedicated training. We find that the resulting directions share a common authorship backbone while conflicting on aspect-specific residuals that carry genuine stylistic signal, explaining why naive aggregation fails. We operationalize this in Aspect-Aware Activation Steering (A3S), a training-free framework that merges per-aspect contrastive directions with interference-aware aggregation and tunes steering strength per instance. A3S improves authorship style transfer where it is genuinely multi-aspect, outperforms a trained baseline in preference evaluations on out-of-domain benchmarks, and keeps target-exemplar overlap consistently low.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。