arXiv:2507.18033cs.ROcs.AI2025-07被引 6

用多模态大模型让机器人理解自然语言指令,自主规划路径导航。

OpenNav: Open-World Navigation with Multimodal Large Language Models

  • 通过多模态大模型解析语言指令,生成鸟瞰图价值地图。
  • 零样本下在真实场景中完成多样导航任务,抗误检与歧义。
  • 适合做开放世界导航的科研与工程人员参考。

预训练的大语言模型(LLMs)展现出强大的常识推理能力,为机器人导航与规划任务带来新可能。然而,如何在开放世界中将语言描述转化为具体机器人动作,仍面临挑战,尤其在不依赖预设动作基元的情况下。本文提出一种基于多模态大语言模型(MLLMs)的开放世界导航框架,使机器人能够解析复杂自然语言指令,并生成轨迹点序列以完成多样化导航任务。我们发现,MLLMs在处理自由形式语言指令时具备出色的跨模态理解能力,能实现稳健的场景认知。更重要的是,利用其代码生成能力,可与视觉-语言感知模型交互,生成组合式二维鸟瞰图价值地图,有效融合语义知识与空间信息,增强机器人的空间理解。我们基于大规模自动驾驶数据集(AVDs)验证了该框架在户外导航中的零样本性能,证明其能执行多种自由形式语言指令,同时对目标检测错误和语言歧义具有鲁棒性。此外,在室内与室外场景中对Husky机器人进行实测,进一步验证系统在真实环境下的有效性与适用性。补充视频见:https://trailab.github.io/OpenNav-website/

原文摘要 · Abstract (English)

Pre-trained large language models (LLMs) have demonstrated strong common-sense reasoning abilities, making them promising for robotic navigation and planning tasks. However, despite recent progress, bridging the gap between language descriptions and actual robot actions in the open-world, beyond merely invoking limited predefined motion primitives, remains an open challenge. In this work, we aim to enable robots to interpret and decompose complex language instructions, ultimately synthesizing a sequence of trajectory points to complete diverse navigation tasks given open-set instructions and open-set objects. We observe that multi-modal large language models (MLLMs) exhibit strong cross-modal understanding when processing free-form language instructions, demonstrating robust scene comprehension. More importantly, leveraging their code-generation capability, MLLMs can interact with vision-language perception models to generate compositional 2D bird-eye-view value maps, effectively integrating semantic knowledge from MLLMs with spatial information from maps to reinforce the robot's spatial understanding. To further validate our approach, we effectively leverage large-scale autonomous vehicle datasets (AVDs) to validate our proposed zero-shot vision-language navigation framework in outdoor navigation tasks, demonstrating its capability to execute a diverse range of free-form natural language navigation instructions while maintaining robustness against object detection errors and linguistic ambiguities. Furthermore, we validate our system on a Husky robot in both indoor and outdoor scenes, demonstrating its real-world robustness and applicability. Supplementary videos are available at https://trailab.github.io/OpenNav-website/

机器人导航多模态模型语言理解开放世界

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。