arXiv:2601.22350cs.LGcs.AI2026-01被引 1

学习策略的统一表示,实现测试时无需训练的行为精准调控。

Learning Policy Representations for Steerable Behavior Synthesis

  • 用状态-动作特征期望建模策略表示,基于占用度量统一表达不同策略。
  • 通过集合架构和变分生成,实现对多种策略及其价值函数的联合编码。
  • 在潜在空间中直接梯度优化,可零样本满足新价值约束条件。

给定一个马尔可夫决策过程(MDP),我们旨在学习一系列策略的统一表示,以支持测试阶段的行为调控。由于MDP的策略由其占用度量唯一决定,我们提出将策略表示建模为相对于占用度量的状态-动作特征映射的期望。我们证明,通过集合架构可以一致地近似一组策略的这些表示。模型将一组状态-动作样本编码为潜在嵌入,从中解码出对应多个奖励的策略及其价值函数。采用变分生成方法诱导平滑的潜在空间,并通过对比学习进一步调整,使潜在距离与价值函数差异对齐。该几何结构允许在潜在空间中进行直接的梯度优化。利用此能力,我们解决了一个新颖的行为合成任务:在不进行额外训练的情况下,将策略调控以满足先前未见的价值函数约束。

原文摘要 · Abstract (English)

Given a Markov decision process (MDP), we seek to learn representations for a range of policies to facilitate behavior steering at test time. As policies of an MDP are uniquely determined by their occupancy measures, we propose modeling policy representations as expectations of state-action feature maps with respect to occupancy measures. We show that these representations can be approximated uniformly for a range of policies using a set-based architecture. Our model encodes a set of state-action samples into a latent embedding, from which we decode both the policy and its value functions corresponding to multiple rewards. We use variational generative approach to induce a smooth latent space, and further shape it with contrastive learning so that latent distances align with differences in value functions. This geometry permits gradient-based optimization directly in the latent space. Leveraging this capability, we solve a novel behavior synthesis task, where policies are steered to satisfy previously unseen value function constraints without additional training.

策略表示行为调控潜在空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。