用连续潜变量控制冻结策略,让模型按偏好生成更好行为。
Preference Goal Tuning: Post-Training as Latent Control for Frozen Policies
- 将目标嵌入作为连续控制变量,优化潜空间目标来调节行为。
- 在Minecraft 17个任务上,平均提升72%~81.6%,优于专家提示。
- 保持策略冻结,适合数据少或需鲁棒泛化的场景。
目标条件策略可根据指定目标执行多样化行为,但下游性能常高度依赖指令或提示词。为克服离散文本提示的局限,我们将后训练适配建模为潜空间控制问题,其中目标嵌入作为连续控制变量以调节冻结策略的行为。提出偏好目标调优(PGT)框架,通过轨迹级偏好目标优化该潜空间控制变量,使生成轨迹分布与任务偏好对齐。与更新策略参数的标准微调不同,PGT保持策略冻结,仅更新潜空间目标。该方法本质上是在搜索最优条件输入,以最大化期望行为的可能性并抑制不良行为。我们在Minecraft SkillForge基准上评估了PGT,覆盖17个任务。仅用少量数据,便在两个基础策略上分别实现72.0%和81.6%的平均相对提升,持续优于专家设计的提示。关键在于,通过解耦任务对齐(潜空间目标)与物理动态(冻结策略),PGT在分布外设置下超越全量微调13.4%,展现出更强的鲁棒性和泛化能力。
原文摘要 · Abstract (English)
Goal-conditioned policies enable decision-making models to execute diverse behaviors based on specified goals, yet their downstream performance is often highly sensitive to the choice of instructions or prompts. To bypass the limitations of discrete text prompts, we formulate post-training adaptation as a latent control problem, where the goal embedding serves as a continuous control variable to modulate the behavior of a frozen policy. We propose Preference Goal Tuning (PGT), a framework that optimizes this latent control variable to align the induced trajectory distribution with task preferences. Unlike standard fine-tuning that updates policy parameters, PGT keeps the policy frozen and updates only the latent goal using a trajectory-level preference objective. This approach essentially searches for the optimal conditioning input that maximizes the likelihood of preferred behaviors while suppressing undesirable ones. We evaluate PGT on the Minecraft SkillForge benchmark across 17 tasks. With minimal data, PGT achieves average relative improvements of 72.0\% and 81.6\% on two foundation policies, consistently outperforming expert-crafted prompts. Crucially, by decoupling task alignment (latent goal) from physical dynamics (frozen policy), PGT surpasses full fine-tuning by 13.4\% in out-of-distribution settings, demonstrating superior robustness and generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。