arXiv:2411.12982cs.RO2024-11

用接触引导生成机器人操作轨迹,提升复杂任务的可控性与成功率。

Hierarchical Diffusion Policy: manipulation trajectory generation via contact guidance

  • 分层扩散策略:高层预测接触点,低层生成动作序列。
  • 6项任务平均提升20.8%,在接触丰富的任务中表现更优。
  • 支持真实世界刚性与柔性物体操作,可解释性强。

基于去噪扩散过程的机器人决策日益成为研究热点,但端到端策略在富含接触的任务中表现不佳且可控性差。本文提出分层扩散策略(HDP),一种利用目标接触引导机器人轨迹生成的模仿学习方法。该策略分为两层:高层基于3D信息预测下一步操作的接触点;低层基于观测与接触的隐变量预测动作序列。两层均以条件去噪扩散过程建模,并结合行为克隆与Q-learning优化低层策略,实现对接触点的精准引导。在6个不同任务上的基准测试显示,HDP显著优于现有最先进方法Diffusion Policy,平均性能提升20.8%。接触引导带来显著优势:性能提升、可解释性增强、可控性更强,尤其在接触密集任务中表现突出。为进一步释放HDP潜力,本文提出三项关键技术:快照梯度优化(提升训练效率)、3D条件化(增强空间感知)、提示引导(增强可控性)。最终,真实世界实验验证了HDP能有效处理刚性与柔性物体。

原文摘要 · Abstract (English)

Decision-making in robotics using denoising diffusion processes has increasingly become a hot research topic, but end-to-end policies perform poorly in tasks with rich contact and have limited controllability. This paper proposes Hierarchical Diffusion Policy (HDP), a new imitation learning method of using objective contacts to guide the generation of robot trajectories. The policy is divided into two layers: the high-level policy predicts the contact for the robot's next object manipulation based on 3D information, while the low-level policy predicts the action sequence toward the high-level contact based on the latent variables of observation and contact. We represent both level policies as conditional denoising diffusion processes, and combine behavioral cloning and Q-learning to optimize the low level policy for accurately guiding actions towards contact. We benchmark Hierarchical Diffusion Policy across 6 different tasks and find that it significantly outperforms the existing state of-the-art imitation learning method Diffusion Policy with an average improvement of 20.8%. We find that contact guidance yields significant improvements, including superior performance, greater interpretability, and stronger controllability, especially on contact-rich tasks. To further unlock the potential of HDP, this paper proposes a set of key technical contributions including snapshot gradient optimization, 3D conditioning, and prompt guidance, which improve the policy's optimization efficiency, spatial awareness, and controllability respectively. Finally, real world experiments verify that HDP can handle both rigid and deformable objects.

机器人操作扩散模型接触引导分层策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。