arXiv:2411.15131cs.ROcs.CV2024-11ICRA被引 24

让四足机器人在真实环境中完成长时间复杂操作任务

WildLMa: Long Horizon Loco-Manipulation in the Wild

  • 用视觉-语言模型实现语言指令驱动的技能学习
  • 仅用几十次示范就达到比强化学习更高的抓取成功率
  • 适合需要长时序、跨场景操作的机器人应用

现实世界中的移动操作要求机器人具备跨物体配置的泛化能力、在多样环境中执行长时序任务的能力,以及超越抓放的复杂操作能力。四足带机械臂机器人有望拓展工作范围并实现稳健移动,但现有研究未充分探索此能力。本文提出WildLMa,包含三个组件:(1) 基于VR的全身遥操作与可通行性适配的低层控制器;(2) WildLMa-Skill——通过模仿学习或启发式方法获取的通用视觉-运动技能库;(3) WildLMa-Planner——让大语言模型能协调技能完成长时序任务的接口。我们证明高质量训练数据的重要性,仅用数十次示范即实现高于现有强化学习基线的抓取成功率。WildLMa利用CLIP实现语言条件下的模仿学习,实证可泛化至训练中未见物体。除大量定量评估外,还展示了实际应用效果,如在校园走廊或户外地形清理垃圾、操作活动物体、整理书架物品。

原文摘要 · Abstract (English)

'In-the-wild' mobile manipulation aims to deploy robots in diverse real-world environments, which requires the robot to (1) have skills that generalize across object configurations; (2) be capable of long-horizon task execution in diverse environments; and (3) perform complex manipulation beyond pick-and-place. Quadruped robots with manipulators hold promise for extending the workspace and enabling robust locomotion, but existing results do not investigate such a capability. This paper proposes WildLMa with three components to address these issues: (1) adaptation of learned low-level controller for VR-enabled whole-body teleoperation and traversability; (2) WildLMa-Skill -- a library of generalizable visuomotor skills acquired via imitation learning or heuristics and (3) WildLMa-Planner -- an interface of learned skills that allow LLM planners to coordinate skills for long-horizon tasks. We demonstrate the importance of high-quality training data by achieving higher grasping success rate over existing RL baselines using only tens of demonstrations. WildLMa exploits CLIP for language-conditioned imitation learning that empirically generalizes to objects unseen in training demonstrations. Besides extensive quantitative evaluation, we qualitatively demonstrate practical robot applications, such as cleaning up trash in university hallways or outdoor terrains, operating articulated objects, and rearranging items on a bookshelf.

移动操作四足机器人长时序任务语言控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。