用扩散策略提升机器人操作中的不确定性估计,减少错误并加快学习。
Diff-DAgger: Uncertainty Estimation with Diffusion Policy for Robotic Manipulation
- 结合扩散策略的训练目标,改进交互式模仿学习中的不确定性判断。
- 在堆叠、推动等任务中,失败预测准确率提升39.0%,任务完成率提高20.6%。
- 适合希望高效使用高表达力策略进行机器人交互学习的研究者。
最近,扩散策略在处理机器人操作中的多模态任务上表现出色。然而,其在分布外情况下的失败问题依然存在,主要源于误差累积和泛化能力有限。一种解决方法是机器人引导的DAgger,通过机器人主动向专家求助来实现交互式模仿学习。尽管机器人引导DAgger具有大规模学习潜力,但现有方法如集成-DAgger在高表达力策略下表现不佳:常将策略分歧误判为不确定性。为此,我们提出Diff-DAgger,一种高效的机器人引导DAgger算法,利用扩散策略的训练目标。我们在堆叠、推动和插拔等不同机器人任务上评估了该方法,结果显示,任务失败预测准确率提升39.0%,任务完成率提高20.6%,壁钟时间减少7.8倍。我们希望这项工作能为将表达力强但数据需求高的策略有效融入交互式机器人学习提供新路径。项目网站:https://diffdagger.github.io。
原文摘要 · Abstract (English)
Recently, diffusion policy has shown impressive results in handling multi-modal tasks in robotic manipulation. However, it has fundamental limitations in out-of-distribution failures that persist due to compounding errors and its limited capability to extrapolate. One way to address these limitations is robot-gated DAgger, an interactive imitation learning with a robot query system to actively seek expert help during policy rollout. While robot-gated DAgger has high potential for learning at scale, existing methods like Ensemble-DAgger struggle with highly expressive policies: They often misinterpret policy disagreements as uncertainty at multi-modal decision points. To address this problem, we introduce Diff-DAgger, an efficient robot-gated DAgger algorithm that leverages the training objective of diffusion policy. We evaluate Diff-DAgger across different robot tasks including stacking, pushing, and plugging, and show that Diff-DAgger improves the task failure prediction by 39.0%, the task completion rate by 20.6%, and reduces the wall-clock time by a factor of 7.8. We hope that this work opens up a path for efficiently incorporating expressive yet data-hungry policies into interactive robot learning settings. The project website is available at: https://diffdagger.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。