arXiv:2607.01802cs.CL2026-07

研究发现,控制生成的引导向量在实际应用中存在显著局限。

On the Limits of Steering Vectors for Preference-Aligned Generation

论文配图:On the Limits of Steering Vectors for Preference-Aligned Generation
图 1 · 摘自论文原文
  • 通过提取不同风格的引导向量,测试其在文本生成中的表现
  • 向量迁移时效果下降,多特征组合会严重降低表达能力
  • 适合关注可控生成可靠性的研究人员参考

引导向量作为一种可解释、无需训练的文本生成控制方法,近年来受到广泛关注。然而其实际泛化能力仍不明确。本文基于PLUME写作个性化基准,针对多种偏好提取引导向量,并在Qwen2.5-7B-Instruct和Llama3.1-8B-Instruct两个开源模型上,评估其在摘要生成与邮件写作任务中的表现。结果表明,引导向量在不同特质上的有效性差异显著;从正负样本中提取的向量在迁移到下游个性化任务时效果会下降;同时,多种多向量组合方法均表现出随向量数量增加而明显降低的特质表达力,且存在连贯性与表达性之间的权衡,需针对具体场景调参。综合来看,引导向量作为通用偏好对齐工具存在实质性限制。

原文摘要 · Abstract (English)

Steering vectors have emerged as a promising approach to controlled text generation, offering interpretable, training-free mechanisms for shaping model outputs. However, their practical generality remains poorly understood. We study the limits of steering vector generalization along three dimensions: trait expressibility, task transfer, and multi-trait composition. Using the PLUME writing personalization benchmark, we extract steering vectors for a range of preferences and evaluate them on summarization and email-writing tasks across two open-source models (Qwen2.5-7B-Instruct and Llama3.1-8B-Instruct). We find that steering effectiveness varies substantially across traits. We further show that steering effectiveness can degrade when vectors extracted from positive and negative style examples are transferred to downstream writing personalization tasks. Finally, we compare common methods for composing multiple steering vectors and find that all methods suffer significant drops in trait expression as more vectors are added, with a tradeoff between coherence and expressibility that requires per-setting hyperparameter tuning. Taken together, our results suggest that steering vectors face meaningful limits as a general-purpose tool for preference alignment.

文本生成可控生成引导向量偏好对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。