基于最优传输的一步生成策略,提升机器人操作任务表现
Drifting Field Policy: A One-Step Generative Policy via Wasserstein Gradient Flow
- 通过反向KL-Wasserstein梯度流实现概率空间中的策略更新
- 在Robomimic和OGBench上优于基于ODE的策略,达成顶尖性能
- 适合追求高效推理与稳定策略优化的研究者与工程师
我们提出漂移场策略(DFP),一种基于漂移模型范式的非微分方程(non-ODE)一步生成策略。将策略更新建模为向软目标策略的反向KL Wasserstein-2梯度流,使每次更新对应于概率空间中的梯度步。该梯度被分解为向高动作价值区域上升以及与锚定策略进行得分匹配(作为信任区域)。我们进一步推导出原本难以处理的更新损失的简洁可计算近似,类似于在顶级批评者选择的动作上进行行为克隆。实验表明,这一机制对非ODE参数化的漂移主干网络尤为有益。通过一步推理,DFP在Robomimic和OGBench多个操作任务上达到当前最优性能,超越基于ODE的策略。
原文摘要 · Abstract (English)
We propose Drifting Field Policy (DFP), a non-ODE one-step generative policy built on the drifting model paradigm. We frame the policy update as a reverse-KL Wasserstein-2 gradient flow toward a soft target policy, so that each DFP update corresponds to a gradient step in probability space. By construction, this gradient is decomposed into an ascent toward higher action-value regions and a score matching with the anchor policy as a trust region. We further derive a simple, tractable surrogate of the otherwise intractable update loss, akin to behavior cloning on top-K critic-selected actions. We find empirically that this mechanism uniquely benefits the drifting backbone owing to its non-ODE parameterization. With one-step inference, DFP achieves state-of-the-art performance on several manipulation tasks across Robomimic and OGBench, outperforming ODE-based policies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。