arXiv:2606.01151cs.LG2026-06中稿 · ICML被引 1

用噪声扰动优化生成策略,提升模仿学习的样本效率与成功率。

Lagrangian Perturbation Diffusion Steering: Latent Reinforcement Learning for Generative Policies

论文配图:Lagrangian Perturbation Diffusion Steering: Latent Reinforcement Learning for Generative Policies
图 1 · 摘自论文原文
  • 在解码前学习紧凑的噪声空间扰动,轻量微调冻结的生成策略。
  • 在多个基准上实现最高25%的回报提升,同时保持更高动作熵。
  • 适用于仿真与物理机器人,兼容扩散模型和视觉-语言-动作大模型。

高容量生成策略的行为克隆能实现强模仿性能,但常受限于示范数据覆盖不足和分布偏移。直接强化学习微调可提升表现,但更新大型动作解码器往往不稳定且样本效率低。我们提出拉格朗日扰动扩散引导(LP-DS),一种轻量级适应方法:通过在解码前学习紧凑的噪声空间扰动来改进冻结的生成策略。LP-DS 使用拉格朗日信任域目标优化该扰动,在提升下游价值的同时约束对潜在先验的偏离。在 RoboMimic 操控、OpenAI Gym 步行与 Adroit 灵巧操控基准上,LP-DS 提升了样本效率、成功率与回报,且动作空间熵高于无约束噪声空间引导。回报提升最高达25%。额外实验表明,该方法不限于紧凑扩散策略或仿真环境,同样适用于流匹配主干、大型视觉-语言-动作模型及真实 Franka 机械臂部署。

原文摘要 · Abstract (English)

Behavior cloning with high-capacity generative policies achieves strong imitation performance, but is often limited by demonstration coverage and distribution shift. Direct reinforcement learning fine-tuning can improve performance, but updating large action decoders is frequently unstable and sample inefficient. We propose Lagrangian Perturbation Diffusion Steering (LP-DS), a lightweight adaptation method that improves a frozen generative policy by learning a compact noise-space perturbation before decoding. LP-DS optimizes this perturbation with a Lagrangian trust-region objective, improving downstream value while constraining deviation from the latent prior. Across RoboMimic manipulation, OpenAI Gym locomotion, and Adroit dexterous manipulation benchmarks, LP-DS improves sample efficiency, success, and return while maintaining higher action-space entropy than unconstrained noise-space steering, with return improvements of up to 25% over prior baselines. Additional evaluations with flow-matching backbones, a large vision-language-action model, and physical Franka deployment show that LP-DS is not limited to compact diffusion policies or simulated benchmarks. Project page: https://sites.google.com/view/lp-ds/home.

生成策略强化学习扩散模型机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。