arXiv:2412.04453cs.ROcs.CV2024-12被引 263

让机器人听懂指令并自主走复杂地形,通过分步规划提升导航准确性。

NaVILA: Legged Robot Vision-Language-Action Model for Navigation

  • 分两层设计:先生成带距离的中层指令,再由视觉强化学习执行
  • 在新构建的IsaacLab基准上,导航成功率显著高于现有方法
  • 支持真实机器人实验,适用于复杂环境中的语言控制导航

本文针对腿式机器人进行视觉-语言导航,不仅为人类提供灵活指令方式,还能使机器人在更复杂、拥挤的场景中导航。然而,将语言指令直接转化为低层腿部关节动作极具挑战。我们提出NaVILA,一种两级框架,统一视觉-语言-动作模型(VLA)与运动技能。不同于直接从VLA预测底层动作,NaVILA首先生成带有空间信息的中层动作(如“向前移动75cm”),作为视觉运动强化学习策略的输入以执行。NaVILA在现有基准上显著优于先前方法。在新构建的IsaacLab基准上进一步验证,该基准包含更逼真的场景、底层控制及真实机器人实验。更多结果见https://navila-bot.github.io/

原文摘要 · Abstract (English)

This paper proposes to solve the problem of Vision-and-Language Navigation with legged robots, which not only provides a flexible way for humans to command but also allows the robot to navigate through more challenging and cluttered scenes. However, it is non-trivial to translate human language instructions all the way to low-level leg joint actions. We propose NaVILA, a 2-level framework that unifies a Vision-Language-Action model (VLA) with locomotion skills. Instead of directly predicting low-level actions from VLA, NaVILA first generates mid-level actions with spatial information in the form of language, (e.g., "moving forward 75cm"), which serves as an input for a visual locomotion RL policy for execution. NaVILA substantially improves previous approaches on existing benchmarks. The same advantages are demonstrated in our newly developed benchmarks with IsaacLab, featuring more realistic scenes, low-level controls, and real-world robot experiments. We show more results at https://navila-bot.github.io/

机器人导航多模态强化学习语言指令

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。