arXiv:2503.13446cs.ROcs.CV2025-03CVPR被引 56

将固定基机器人模型迁移到移动机器人,实现零样本自适应

MoManipVLA: Transferring Vision-language-action Models for General Mobile Manipulation

  • 用预训练视觉-语言-动作模型生成末端执行器路径点
  • 设计双层优化框架,提升移动基座与机械臂轨迹可行性
  • 实测成功率超当前最优4.2%,部署仅需50次训练

移动操作是机器人在日常生活中协助人类完成多样化任务和环境适应的核心挑战。传统方法因缺乏大规模训练而泛化能力差。尽管近期视觉-语言-动作(VLA)模型展现出强大泛化能力,但其针对固定基操作任务设计。为此,我们提出高效策略迁移框架MoManipVLA,将预训练的固定基VLA模型迁移到移动操作中,从而实现跨任务与环境的高泛化性。具体而言,利用预训练VLA模型生成具有强泛化能力的末端执行器路径点,设计移动基座与机械臂的运动规划目标,最大化轨迹物理可行性。提出高效的双层目标优化框架:上层优化预测基座移动路径以扩展操作器策略空间,下层优化选择最优末端执行器轨迹完成任务。该方法可零样本调整机器人基座位置,使原固定基模型生成的路径可行。在OVMM数据集及真实世界中的大量实验表明,MoManipVLA成功率比当前最优方法高4.2%,且由于预训练模型强泛化性,真实部署仅需50次训练成本。

原文摘要 · Abstract (English)

Mobile manipulation is the fundamental challenge for robotics to assist humans with diverse tasks and environments in everyday life. However, conventional mobile manipulation approaches often struggle to generalize across different tasks and environments because of the lack of large-scale training. In contrast, recent advances in vision-language-action (VLA) models have shown impressive generalization capabilities, but these foundation models are developed for fixed-base manipulation tasks. Therefore, we propose an efficient policy adaptation framework named MoManipVLA to transfer pre-trained VLA models of fix-base manipulation to mobile manipulation, so that high generalization ability across tasks and environments can be achieved in mobile manipulation policy. Specifically, we utilize pre-trained VLA models to generate waypoints of the end-effector with high generalization ability. We design motion planning objectives for the mobile base and the robot arm, which aim at maximizing the physical feasibility of the trajectory. Finally, we present an efficient bi-level objective optimization framework for trajectory generation, where the upper-level optimization predicts waypoints for base movement to enhance the manipulator policy space, and the lower-level optimization selects the optimal end-effector trajectory to complete the manipulation task. In this way, MoManipVLA can adjust the position of the robot base in a zero-shot manner, thus making the waypoints predicted from the fixed-base VLA models feasible. Extensive experimental results on OVMM and the real world demonstrate that MoManipVLA achieves a 4.2% higher success rate than the state-of-the-art mobile manipulation, and only requires 50 training cost for real world deployment due to the strong generalization ability in the pre-trained VLA models.

移动操作视觉语言动作零样本迁移双层优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。