让机器人零样本完成移动操作,自动找合适位置动手
MoTo: A Zero-shot Plug-in Interaction-aware Navigation for General Mobile Manipulation
- 用视觉语言模型识别物体和机械臂关键点,实现零样本交互感知
- 生成机器人停靠点,使固定基座模型可直接用于移动操作
- 无需专家数据,在真实场景中成功率提升16.67%
移动操作是机器人领域核心挑战,使机器人能在多样任务与动态环境中协助人类。传统方法因缺乏大规模训练而泛化能力差。近期操纵基础模型在固定基座任务上表现优异,但受限于固定设置。为此,我们提出可插拔模块MoTo,能与任意现成操纵基础模型结合,赋予其移动操作能力。具体地,设计了交互感知导航策略,生成机器人停靠点以支持通用移动操作。为实现零样本能力,提出基于多视角一致性的视觉语言模型交互关键点框架,用于目标物体和机械臂跟随指令,使固定基座模型可直接使用。进一步设计移动基座与机械臂的运动规划目标,最小化两点间距离并保证轨迹物理可行性。通过此方式,MoTo引导机器人移动至可执行固定基座操作的位置,并利用VLM生成与轨迹优化实现零样本移动操作,无需任何移动操作专家数据。在OVMM和真实世界实验中,成功率达2.68%和16.67%高于当前最优方法,且无需额外训练数据。
原文摘要 · Abstract (English)
Mobile manipulation stands as a core challenge in robotics, enabling robots to assist humans across varied tasks and dynamic daily environments. Conventional mobile manipulation approaches often struggle to generalize across different tasks and environments due to the lack of large-scale training. However, recent advances in manipulation foundation models demonstrate impressive generalization capability on a wide range of fixed-base manipulation tasks, which are still limited to a fixed setting. Therefore, we devise a plug-in module named MoTo, which can be combined with any off-the-shelf manipulation foundation model to empower them with mobile manipulation ability. Specifically, we propose an interaction-aware navigation policy to generate robot docking points for generalized mobile manipulation. To enable zero-shot ability, we propose an interaction keypoints framework via vision-language models (VLM) under multi-view consistency for both target object and robotic arm following instructions, where fixed-base manipulation foundation models can be employed. We further propose motion planning objectives for the mobile base and robot arm, which minimize the distance between the two keypoints and maintain the physical feasibility of trajectories. In this way, MoTo guides the robot to move to the docking points where fixed-base manipulation can be successfully performed, and leverages VLM generation and trajectory optimization to achieve mobile manipulation in a zero-shot manner, without any requirement on mobile manipulation expert data. Extensive experimental results on OVMM and real-world demonstrate that MoTo achieves success rates of 2.68% and 16.67% higher than the state-of-the-art mobile manipulation methods, respectively, without requiring additional training data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。