用局部策略实现无需训练的长序列机器人操作,真实场景表现卓越。
Local Policies Enable Zero-shot Long-horizon Manipulation
- 采用局部策略,对位姿和场景变化具有不变性。
- 仿真中零样本任务成功率97%,真实场景解决8阶段复杂任务。
- 适合需要快速部署、泛化能力强的机器人应用。
机器人操作的模拟到现实(sim2real)迁移因复杂接触建模和真实任务分布生成困难而面临挑战。为解决后者问题,我们提出ManipGen,引入一类新型策略——局部策略。局部性带来多种优势,包括对机器人与物体绝对位姿、技能顺序及全局场景配置的不变性。我们将此类策略与视觉、语言和运动规划的基础模型结合,在仿真中实现了对Robosuite基准任务的当前最佳零样本性能(97%)。将局部策略从仿真迁移到现实后,其可在存在显著位姿、物体和场景配置差异的情况下,成功完成多达8个阶段的未见长序列操作任务。在50个真实世界操作任务上,ManipGen相较于SayCan、OpenVLA、LLMTrajGen和VoxPoser分别提升36%、76%、62%和60%。视频结果见https://mihdalal.github.io/manipgen/
原文摘要 · Abstract (English)
Sim2real for robotic manipulation is difficult due to the challenges of simulating complex contacts and generating realistic task distributions. To tackle the latter problem, we introduce ManipGen, which leverages a new class of policies for sim2real transfer: local policies. Locality enables a variety of appealing properties including invariances to absolute robot and object pose, skill ordering, and global scene configuration. We combine these policies with foundation models for vision, language and motion planning and demonstrate SOTA zero-shot performance of our method to Robosuite benchmark tasks in simulation (97%). We transfer our local policies from simulation to reality and observe they can solve unseen long-horizon manipulation tasks with up to 8 stages with significant pose, object and scene configuration variation. ManipGen outperforms SOTA approaches such as SayCan, OpenVLA, LLMTrajGen and VoxPoser across 50 real-world manipulation tasks by 36%, 76%, 62% and 60% respectively. Video results at https://mihdalal.github.io/manipgen/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。