arXiv:2602.10399cs.RO2026-02被引 1

让机器人听懂指令,实时调整走路方式。

LocoVLM: Grounding Vision and Language for Adapting Versatile Legged Locomotion Policies

  • 用大模型生成指令对应的技能库,结合视觉语义定位环境。
  • 实现87%指令跟随准确率,无需联网调用云端模型。
  • 可灵活切换多种行走风格,适合复杂场景的机器人控制。

当前足式运动学习仍依赖环境的几何表示,限制了机器人对高层次语义(如人类指令)的响应能力。为此,我们提出一种新方法,将基础模型中的常识推理能力融入足式运动适应过程。具体而言,利用预训练大语言模型生成面向机器人的指令-技能数据库;通过预训练视觉语言模型提取环境高层语义并将其在技能库中进行定位,从而实现实时技能建议。为支持多样化技能控制,我们训练了一个风格条件策略,能以高保真度生成符合指定风格的多样化、鲁棒的运动技能。据我们所知,这是首个在无需在线查询云端基础模型的前提下,实现基于环境语义和指令的实时足式运动适应的工作,指令遵循准确率达87%。

原文摘要 · Abstract (English)

Recent advances in legged locomotion learning are still dominated by the utilization of geometric representations of the environment, limiting the robot's capability to respond to higher-level semantics such as human instructions. To address this limitation, we propose a novel approach that integrates high-level commonsense reasoning from foundation models into the process of legged locomotion adaptation. Specifically, our method utilizes a pre-trained large language model to synthesize an instruction-grounded skill database tailored for legged robots. A pre-trained vision-language model is employed to extract high-level environmental semantics and ground them within the skill database, enabling real-time skill advisories for the robot. To facilitate versatile skill control, we train a style-conditioned policy capable of generating diverse and robust locomotion skills with high fidelity to specified styles. To the best of our knowledge, this is the first work to demonstrate real-time adaptation of legged locomotion using high-level reasoning from environmental semantics and instructions with instruction-following accuracy of up to 87% without the need for online query to on-the-cloud foundation models.

足式机器人指令理解视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。