让机械臂具备3D前瞻能力,提升复杂抓取成功率
3D Dynamics-Aware Manipulation: Endowing Manipulation Policies with 3D Foresight
- 通过自监督学习构建3D世界模型,融合深度估计与运动预测
- 在仿真和真实场景中,任务成功率显著提升,推理速度不变
- 适合需要精确空间操作的机器人抓取、装配任务
将世界建模融入抓取策略学习,已推动抓取性能边界。但现有方法仅建模2D视觉动态,在涉及明显深度运动的任务中表现不足。为此,我们提出一种3D动态感知抓取框架,无缝集成3D世界建模与策略学习。框架内引入三项自监督学习任务:当前深度估计、未来RGB-D预测和3D运动流预测,三者互补,赋予策略模型3D前瞻性。大量仿真与真实世界实验表明,3D前瞻性可显著提升抓取策略性能,且不牺牲推理速度。代码已开源:https://github.com/Stardust-hyx/3D-Foresight。
原文摘要 · Abstract (English)
The incorporation of world modeling into manipulation policy learning has pushed the boundary of manipulation performance. However, existing efforts simply model the 2D visual dynamics, which is insufficient for robust manipulation when target tasks involve prominent depth-wise movement. To address this, we present a 3D dynamics-aware manipulation framework that seamlessly integrates 3D world modeling and policy learning. Three self-supervised learning tasks (current depth estimation, future RGB-D prediction, 3D flow prediction) are introduced within our framework, which complement each other and endow the policy model with 3D foresight. Extensive experiments on simulation and the real world show that 3D foresight can greatly boost the performance of manipulation policies without sacrificing inference speed. Code is available at https://github.com/Stardust-hyx/3D-Foresight.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。