arXiv:2411.14913cs.RO2024-11中稿 · ICRA被引 6

用扩散模型提升非抓取操作的探索能力,显著提高真实场景成功率。

Enhancing Exploration with Diffusion Policies in Hybrid Off-Policy RL: Application to Non-Prehensile Manipulation

  • 将连续动作建模为扩散模型,结合最大熵强化学习统一处理离散与连续动作。
  • 在真实6D姿态对齐任务中,成功率从53%提升至72%。
  • 适合需要多样化行为策略的机器人操控场景,尤其关注零样本迁移应用。

非抓取操作的多样化策略学习对于提升技能迁移能力和分布外场景泛化至关重要。本文提出一种混合框架下的双重探索增强方法,同时处理离散和连续动作空间。首先,将连续运动参数策略建模为扩散模型;其次,将其融入最大熵强化学习框架,统一优化离散与连续部分。离散动作(如接触点选择)通过Q值函数最大化进行优化,连续部分则由基于扩散的策略引导。该混合方法具有理论依据,最大熵项通过结构化变分推断得到下界。我们提出混合扩散策略算法(HyDo),并在仿真和零样本模拟到现实任务中进行评估。结果表明,HyDo能有效促进多样化行为策略,显著提升各项任务成功率——例如,在真实世界6D姿态对齐任务中,成功率从53%提升至72%。

原文摘要 · Abstract (English)

Learning diverse policies for non-prehensile manipulation is essential for improving skill transfer and generalization to out-of-distribution scenarios. In this work, we enhance exploration through a two-fold approach within a hybrid framework that tackles both discrete and continuous action spaces. First, we model the continuous motion parameter policy as a diffusion model, and second, we incorporate this into a maximum entropy reinforcement learning framework that unifies both the discrete and continuous components. The discrete action space, such as contact point selection, is optimized through Q-value function maximization, while the continuous part is guided by a diffusion-based policy. This hybrid approach leads to a principled objective, where the maximum entropy term is derived as a lower bound using structured variational inference. We propose the Hybrid Diffusion Policy algorithm (HyDo) and evaluate its performance on both simulation and zero-shot sim2real tasks. Our results show that HyDo encourages more diverse behavior policies, leading to significantly improved success rates across tasks - for example, increasing from 53% to 72% on a real-world 6D pose alignment task. Project page: https://leh2rng.github.io/hydo

强化学习扩散模型机器人操控零样本迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。