测试视觉语言模型在不同移动能力下的室内导航表现
CapNav: Benchmarking Vision Language Models on Capability-conditioned Indoor Navigation
- 设计五类具不同物理能力的智能体,评估其导航适应性
- 涵盖45个真实场景、473项任务,验证模型对空间约束的理解
- 发现顶尖模型在复杂障碍推理上仍严重受限,适合研究具身智能
视觉语言模型(VLMs)在视觉-语言导航(VLN)中取得显著进展,为机器人与人类用户提供新的决策可能。然而,现实导航本质上受制于代理的移动能力。例如,扫地机器人无法走楼梯,而四足机器人可以。我们提出能力条件导航(CapNav)基准,用于评估VLM在特定物理与操作能力条件下,导航复杂室内空间的能力。CapNav定义了五类典型人与机器人代理,分别描述其尺寸、移动能力及环境交互能力。该基准包含45个真实室内场景、473个导航任务和2365个问答对,用于检验VLM是否能根据代理能力正确规划路径。我们评估了13个现代VLM,发现随着移动限制收紧,当前VLM导航性能急剧下降,且即使最先进模型在需空间维度推理的障碍物面前也表现不佳。最后讨论了具备能力意识导航的意义,以及未来提升具身空间推理能力的机会。基准已公开于 https://github.com/makeabilitylab/CapNav
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) have shown remarkable progress in Vision-Language Navigation (VLN), offering new possibilities for navigation decision-making that could benefit both robotic platforms and human users. However, real-world navigation is inherently conditioned by the agent's mobility constraints. For example, a sweeping robot cannot traverse stairs, while a quadruped can. We introduce Capability-Conditioned Navigation (CapNav), a benchmark designed to evaluate how well VLMs can navigate complex indoor spaces given an agent's specific physical and operational capabilities. CapNav defines five representative human and robot agents, each described with physical dimensions, mobility capabilities, and environmental interaction abilities. CapNav provides 45 real-world indoor scenes, 473 navigation tasks, and 2365 QA pairs to test if VLMs can traverse indoor environments based on agent capabilities. We evaluate 13 modern VLMs and find that current VLM's navigation performance drops sharply as mobility constraints tighten, and that even state-of-the-art models struggle with obstacle types that require reasoning on spatial dimensions. We conclude by discussing the implications for capability-aware navigation and the opportunities for advancing embodied spatial reasoning in future VLMs. The benchmark is available at https://github.com/makeabilitylab/CapNav
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。