arXiv:2506.09422cs.ROcs.LG2025-06被引 5

提出时间统一扩散策略,提升机器人操作生成速度与精度。

Time-Unified Diffusion Policy with Action Discrimination for Robotic Manipulation

  • 构建统一时间步的行动去噪速度场,减少训练难度
  • 多视角下成功率82.6%,单视角达83.8%,迭代次数少时优势更明显
  • 适合需要快速精准动作生成的真实场景任务

在复杂场景中,机器人操作依赖生成模型估计多种成功动作的分布。扩散模型因训练鲁棒性优于其他生成模型,在通过成功示范进行模仿学习时表现优异。然而,基于扩散的动作策略通常需多次迭代去噪,难以实现实时响应。此外,现有方法建模随时间变化的动作去噪过程,其时间复杂度增加训练难度并导致动作精度下降。为此,我们提出时间统一扩散策略(TUDP),利用动作识别能力构建统一时间步的去噪过程。一方面,在动作空间构建含额外动作判别信息的时间统一速度场,统一所有时间步的去噪过程,降低策略学习难度并加速动作生成;另一方面,提出动作级训练方法,引入动作判别分支提供额外判别信息,使TUDP隐式学习区分成功动作的能力,提升去噪精度。该方法在RLBench上取得当前最优性能,多视角设置下成功率达82.6%,单视角达83.8%。尤其在较少去噪迭代条件下,成功率提升更为显著。TUDP还能为多种真实任务生成准确动作。

原文摘要 · Abstract (English)

In many complex scenarios, robotic manipulation relies on generative models to estimate the distribution of multiple successful actions. As the diffusion model has better training robustness than other generative models, it performs well in imitation learning through successful robot demonstrations. However, the diffusion-based policy methods typically require significant time to iteratively denoise robot actions, which hinders real-time responses in robotic manipulation. Moreover, existing diffusion policies model a time-varying action denoising process, whose temporal complexity increases the difficulty of model training and leads to suboptimal action accuracy. To generate robot actions efficiently and accurately, we present the Time-Unified Diffusion Policy (TUDP), which utilizes action recognition capabilities to build a time-unified denoising process. On the one hand, we build a time-unified velocity field in action space with additional action discrimination information. By unifying all timesteps of action denoising, our velocity field reduces the difficulty of policy learning and speeds up action generation. On the other hand, we propose an action-wise training method, which introduces an action discrimination branch to supply additional action discrimination information. Through action-wise training, the TUDP implicitly learns the ability to discern successful actions to better denoising accuracy. Our method achieves state-of-the-art performance on RLBench with the highest success rate of 82.6% on a multi-view setup and 83.8% on a single-view setup. In particular, when using fewer denoising iterations, TUDP achieves a more significant improvement in success rate. Additionally, TUDP can produce accurate actions for a wide range of real-world tasks.

机器人操作扩散模型动作生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。