用向量控制生成行为强度,解决模型罕见行为优化难问题
VSPO: Vector-Steered Policy Optimization for Behavioral Control

- 通过可调向量控制生成行为强度,实现行为偏好精准调控
- 在MATH和MMLU-Pro上提升解释专业性与抗误导性,同时保持准确率
- 适合需要精细控制语言模型行为风格的研究者和开发者
现代语言模型常需在主任务准确率之外满足次要行为偏好,如表达详尽度、友善程度或技术深度。实际中,基础模型可能极少甚至完全不具备目标行为。为此,我们提出向量引导策略优化(VSPO),通过关联目标行为的引导向量控制生成结果的行为强度。VSPO通过对GRPO进行修改,以不同引导强度采样轨迹,可视为一种在策略隐空间的自蒸馏过程,使模型内化引导向量。该方法提升了稀有行为的采样频率并丰富轨迹多样性,缓解稀疏奖励问题,并在理论上证明能加速策略优化。在带状抽象下,当引导分布与目标行为对齐时,VSPO的迭代复杂度优于传统奖励塑造的GRPO。我们在MATH和MMLU-Pro等多推理基准上评估了四种目标行为:解释专业性、置信度表达、对误导性上下文的鲁棒性及回答冗长度。结果表明,相比奖励塑造、教师轨迹蒸馏与基于引导的基线,VSPO在增强目标行为控制的同时,保持或提升了任务准确率。
原文摘要 · Abstract (English)
Modern language models often need to optimize a primary accuracy objective while also accommodating secondary behavioral preferences, such as verbosity, agreeableness, or the level of technical expertise in its response. In practice, a base model may exhibit a desired behavior very rarely or not at all. Thus, endowing the model with a target behavior creates a sparse behavioral reward bottleneck. To address such multi-objective problems, we introduce Vector-Steered Policy Optimization (VSPO) which employs a steering vector associated with the target behavior to control the behavior intensity of the generated rollouts. VSPO is obtained by modifying GRPO to sample rollouts with varying steering intensities. This process can be interpreted as an on-policy latent self-distillation procedure where the model internalizes its steering vector. By varying steering intensities, VSPO upsamples rare behaviors and enriches rollout diversity, which alleviates the sparse reward issue and provably accelerates the policy optimization. Through comprehensive theory and experiments, we establish that VSPO has favorable properties compared to vanilla reward shaping and other alternative approaches. Specifically, under a bandit abstraction, VSPO provably achieves better iteration complexity than reward-shaped GRPO when the steering-induced distributions are sufficiently aligned with the target behavior. We evaluate VSPO across multiple reasoning benchmarks, including MATH and MMLU-Pro, for four target behaviors: explanation expertise, confidence expression, robustness to misleading context, and response verbosity. Our results show that VSPO consistently improves the control along target behavior while maintaining or improving task accuracy compared with reward shaping, teacher-trace distillation, and guidance-based baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。