arXiv:2506.15799cs.ROcs.LG2025-06被引 161

用强化学习高效优化扩散策略,无需重训练即可实现机器人自主改进。

Steering Your Diffusion Policy with Latent Space Reinforcement Learning

  • 在扩散策略的噪声空间中直接做强化学习,不修改原始模型参数。
  • 仅需少量真实环境数据(<100次交互)即实现性能显著提升。
  • 适合需要快速适应新任务的机器人系统,尤其擅长通用策略微调。

基于人类示范学习的机器人控制策略在众多现实场景中表现优异。然而,在初始性能不佳的新开放世界场景中,这类行为克隆(BC)策略通常需收集更多人类示范以改进行为,过程昂贵且耗时。相比之下,强化学习(RL)虽可实现自主在线优化,但常因样本需求量大而难以落地。本文提出扩散策略的强化学习引导方法(DSRL):在扩散策略的潜在噪声空间中运行强化学习,实现对BC策略的高效自适应。实验表明,DSRL样本效率高,仅需黑盒访问原始策略,且无需修改基础策略权重,有效避免了微调扩散模型的常见挑战。在仿真基准、真实机器人任务及预训练通用策略适配中均验证了其高效性与实用性。

原文摘要 · Abstract (English)

Robotic control policies learned from human demonstrations have achieved impressive results in many real-world applications. However, in scenarios where initial performance is not satisfactory, as is often the case in novel open-world settings, such behavioral cloning (BC)-learned policies typically require collecting additional human demonstrations to further improve their behavior -- an expensive and time-consuming process. In contrast, reinforcement learning (RL) holds the promise of enabling autonomous online policy improvement, but often falls short of achieving this due to the large number of samples it typically requires. In this work we take steps towards enabling fast autonomous adaptation of BC-trained policies via efficient real-world RL. Focusing in particular on diffusion policies -- a state-of-the-art BC methodology -- we propose diffusion steering via reinforcement learning (DSRL): adapting the BC policy by running RL over its latent-noise space. We show that DSRL is highly sample efficient, requires only black-box access to the BC policy, and enables effective real-world autonomous policy improvement. Furthermore, DSRL avoids many of the challenges associated with finetuning diffusion policies, obviating the need to modify the weights of the base policy at all. We demonstrate DSRL on simulated benchmarks, real-world robotic tasks, and for adapting pretrained generalist policies, illustrating its sample efficiency and effective performance at real-world policy improvement.

扩散模型强化学习机器人控制在线优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。