让人形机器人在大空间中精准移动并操作物体
WholeBodyVLA: Towards Unified Latent VLA for Whole-Body Loco-Manipulation Control
- 用无需动作的视觉语言视频学习全身运动与操作知识
- 在AgiBot X2上比基线提升21.3%,支持复杂任务泛化
- 适合研究人形机器人运动与操作一体化控制的团队
人形机器人需精确运动与灵巧操作以完成复杂任务,但现有方法在运动中缺乏操作感知,限制了工作空间。原因在于:(1) 人形遥操作数据稀缺导致运动-操作知识难获取;(2) 现有强化学习控制器精度与稳定性不足。为此,提出统一潜在学习框架,使视觉-语言-动作(VLA)系统从低成本无动作的自视角视频中学习。同时设计高效人类数据采集流程以扩充数据集。为更精准执行运动指令,构建面向运动-操作的(LMO)强化学习策略,专门优化前进、转向、蹲下等核心动作。基于上述组件,提出WholeBodyVLA,实现人形机器人大空间运动-操作一体化控制。在AgiBot X2上验证,性能优于基线21.3%,展现出强泛化性与高可扩展性。
原文摘要 · Abstract (English)
Humanoid robots require precise locomotion and dexterous manipulation to perform challenging loco-manipulation tasks. Yet existing approaches, modular or end-to-end, are deficient in manipulation-aware locomotion. This confines the robot to a limited workspace, preventing it from performing large-space loco-manipulation. We attribute this to: (1) the challenge of acquiring loco-manipulation knowledge due to the scarcity of humanoid teleoperation data, and (2) the difficulty of faithfully and reliably executing locomotion commands, stemming from the limited precision and stability of existing RL controllers. To acquire richer loco-manipulation knowledge, we propose a unified latent learning framework that enables Vision-Language-Action (VLA) system to learn from low-cost action-free egocentric videos. Moreover, an efficient human data collection pipeline is devised to augment the dataset and scale the benefits. To execute the desired locomotion commands more precisely, we present a loco-manipulation-oriented (LMO) RL policy specifically tailored for accurate and stable core loco-manipulation movements, such as advancing, turning, and squatting. Building on these components, we introduce WholeBodyVLA, a unified framework for humanoid loco-manipulation. To the best of our knowledge, WholeBodyVLA is one of its kind enabling large-space humanoid loco-manipulation. It is verified via comprehensive experiments on the AgiBot X2 humanoid, outperforming prior baseline by 21.3%. It also demonstrates strong generalization and high extensibility across a broad range of tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。