用可交互的度量世界表示,让大模型更懂导航
One Agent to Guide Them All: Empowering MLLMs for Vision-and-Language Navigation via Explicit World Representation
- 将空间感知与语义规划解耦,用度量地图增强大模型推理
- 零样本下在R2R-CE达48.8%成功率,超越现有方法
- 支持跨平台真实机器人部署,适合做通用导航系统
可导航智能体需同时理解高层语义指令与精确空间感知。基于多模态大语言模型(MLLM)的导航系统因其强大泛化能力展现前景,但当前紧耦合设计严重限制性能。本文提出解耦架构,将低层空间状态估计与高层语义规划分离。不同于依赖预设简化文本地图的方法,我们引入可交互的度量世界表示,保持丰富一致信息,使MLLM能在此上进行决策推理。同时引入反事实推理以激发MLLM潜能,度量地图确保动作物理合理性。在模拟与真实环境开展全面实验,方法在零样本条件下达到新纪录:R2R-CE基准48.8%成功率,RxR-CE基准42.2%。进一步验证度量表示的通用性,实现跨不同机器人形态的零样本仿真到现实迁移,包括轮式TurtleBot 4和自研无人机。真实部署结果表明,该解耦框架可作为鲁棒、领域无关的具身视觉-语言导航接口。
原文摘要 · Abstract (English)
A navigable agent needs to understand both high-level semantic instructions and precise spatial perceptions. Building navigation agents centered on Multimodal Large Language Models (MLLMs) demonstrates a promising solution due to their powerful generalization ability. However, the current tightly coupled design dramatically limits system performance. In this work, we propose a decoupled design that separates low-level spatial state estimation from high-level semantic planning. Unlike previous methods that rely on predefined, oversimplified textual maps, we introduce an interactive metric world representation that maintains rich and consistent information, allowing MLLMs to interact with and reason on it for decision-making. Furthermore, counterfactual reasoning is introduced to further elicit MLLMs' capacity, while the metric world representation ensures the physical validity of the produced actions. We conduct comprehensive experiments in both simulated and real-world environments. Our method establishes a new zero-shot state-of-the-art, achieving 48.8\% Success Rate (SR) in R2R-CE and 42.2\% in RxR-CE benchmarks. Furthermore, to validate the versatility of our metric representation, we demonstrate zero-shot sim-to-real transfer across diverse embodiments, including a wheeled TurtleBot 4 and a custom-built aerial drone. These real-world deployments verify that our decoupled framework serves as a robust, domain-invariant interface for embodied Vision-and-Language navigation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。