arXiv:2503.12533cs.ROcs.LG2025-03被引 25

让机器人用语言理解+模块化技能完成复杂任务,实时运行在低成本硬件上。

Being-0: A Humanoid Robotic Agent with Vision-Language Models and Modular Skills

  • 分层架构:大模型负责规划,模块化技能库实现稳定操控。
  • 轻量连接器提升语言指令到动作的转化效率,动态协调行走与抓取。
  • 全尺寸人形机器人实测,可在室内完成多步骤复杂任务。

构建能在真实世界中完成人类级表现的自主人形机器人是该领域的终极目标。近期进展在基础模型(FMs)的高层认知与人形机器人低层技能开发方面取得显著突破。然而,直接组合这些组件常因长时任务中误差累积及各模块延迟差异导致鲁棒性与效率下降。本文提出Being-0,一种融合基础模型与模块化技能库的分层代理框架。基础模型负责指令理解、任务规划与推理等高层认知,技能库则提供稳定的行走与灵巧操作能力。为衔接两层,我们设计了一种由轻量视觉语言模型驱动的新型连接器(Connector),将语言计划转化为可执行技能指令,并动态协调运动与操作以提升任务成功率。除基础模型外,所有组件均可部署于低成本机载计算设备,使Being-0在配备灵巧手和主动视觉的全尺寸人形机器人上实现高效实时运行。大量实验表明,其在大型室内环境中能有效解决需复杂导航与操作的长时程任务。

原文摘要 · Abstract (English)

Building autonomous robotic agents capable of achieving human-level performance in real-world embodied tasks is an ultimate goal in humanoid robot research. Recent advances have made significant progress in high-level cognition with Foundation Models (FMs) and low-level skill development for humanoid robots. However, directly combining these components often results in poor robustness and efficiency due to compounding errors in long-horizon tasks and the varied latency of different modules. We introduce Being-0, a hierarchical agent framework that integrates an FM with a modular skill library. The FM handles high-level cognitive tasks such as instruction understanding, task planning, and reasoning, while the skill library provides stable locomotion and dexterous manipulation for low-level control. To bridge the gap between these levels, we propose a novel Connector module, powered by a lightweight vision-language model (VLM). The Connector enhances the FM's embodied capabilities by translating language-based plans into actionable skill commands and dynamically coordinating locomotion and manipulation to improve task success. With all components, except the FM, deployable on low-cost onboard computation devices, Being-0 achieves efficient, real-time performance on a full-sized humanoid robot equipped with dexterous hands and active vision. Extensive experiments in large indoor environments demonstrate Being-0's effectiveness in solving complex, long-horizon tasks that require challenging navigation and manipulation subtasks. For further details and videos, visit https://beingbeyond.github.io/Being-0.

人形机器人视觉语言模型技能模块化实时控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。