arXiv:2506.01031cs.CV2025-06NeurIPS被引 22

评测大模型在真实环境中的导航能力,发现理解力强的模型执行更准。

NavBench: Probing Multimodal Large Language Models for Embodied Navigation

  • 构建双模块基准:认知理解+分步执行,覆盖3200个问答对和432个场景任务。
  • GPT-4o表现最优,轻量开源模型仅在简单任务中有效,理解力强的模型执行更好。
  • 地图信息提升中等难度任务准确率,但时间进度判断仍是主要短板。

多模态大语言模型(MLLMs)在视觉-语言任务中表现出强大泛化能力,但在具身环境中的理解与行动能力仍待探索。本文提出NavBench,一个用于零样本条件下评估MLLM具身导航能力的基准。该基准包含两部分:(1) 导航理解,通过全局指令对齐、时间进度估计和局部观察能力推理三项认知基础任务,涵盖3,200个问答对;(2) 在72个室内场景中进行432个任务回合的分步执行,按空间、认知与执行复杂度分层。为支持实际部署,引入将模型输出转化为机器人动作的流水线。评估了专有及开源模型,结果表明GPT-4o在各项任务中表现优异,而轻量级开源模型仅在简单情形下有效。结果显示,理解得分高的模型通常执行表现也更好。提供地图上下文可提升决策准确率,尤其在中等难度场景中。然而,大多数模型在时间理解方面存在明显短板,特别是在导航过程中对进展的估计上,这可能构成关键挑战。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have demonstrated strong generalization in vision-language tasks, yet their ability to understand and act within embodied environments remains underexplored. We present NavBench, a benchmark to evaluate the embodied navigation capabilities of MLLMs under zero-shot settings. NavBench consists of two components: (1) navigation comprehension, assessed through three cognitively grounded tasks including global instruction alignment, temporal progress estimation, and local observation-action reasoning, covering 3,200 question-answer pairs; and (2) step-by-step execution in 432 episodes across 72 indoor scenes, stratified by spatial, cognitive, and execution complexity. To support real-world deployment, we introduce a pipeline that converts MLLMs' outputs into robotic actions. We evaluate both proprietary and open-source models, finding that GPT-4o performs well across tasks, while lighter open-source models succeed in simpler cases. Results also show that models with higher comprehension scores tend to achieve better execution performance. Providing map-based context improves decision accuracy, especially in medium-difficulty scenarios. However, most models struggle with temporal understanding, particularly in estimating progress during navigation, which may pose a key challenge.

具身智能多模态模型导航评测大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。