构建首个连续室内外导航基准,测试机器人跨场景能力。
NavVerse: Benchmarking Indoor-to-Outdoor Embodied Navigation in Continuous Robot Simulation

- 设计包含100个室内、50个室外及50个跨场景的物理仿真环境。
- 10,000次任务中,端到端视觉语言模型零样本成功率最高,但安全表现弱于模块化方法。
- 跨场景导航性能显著下降,凸显适应性仍是核心挑战,适合机器人导航研究者参考。
部署于配送、校园及应急响应场景中的机器人常需在单个连续任务中从建筑内导航至街道。现有基准多将室内外导航分开评估,且常忽略机器人执行细节,导致出口寻找、边界穿越、适应性与动力学失败等问题未被充分研究。我们提出NavVerse,一个支持物理仿真的室内外连续导航基准。该基准包含100个室内场景、50个城市室外场景和50个室内外混合场景,覆盖10,000个任务实例,涵盖物体导航、视觉-语言导航和地点导航任务,目标为寻找如餐厅或银行等语义兴趣点。代理通过可执行机器人接口进行评估,使用任务成功率、路径效率与安全性指标。对强化学习、视觉-语言代理(VLA)和模块化基线的零样本实验表明,当前代理距离解决跨上下文导航仍有较大差距:端到端VLA在零样本下表现最佳,而模块化方法安全性最强。PlaceNav实验进一步显示,从室外到室内外混合场景性能明显下降,表明适应性仍是主要瓶颈。
原文摘要 · Abstract (English)
Robots deployed in delivery, campus, and emergency-response settings often need to navigate from buildings to streets within a single continuous episode. Existing benchmarks usually evaluate indoor and outdoor navigation separately, and many abstract away robot execution, leaving exit finding, boundary traversal, adaptation, and kinodynamic failures underexplored. We introduce NavVerse, a physics-enabled benchmark for indoor-to-outdoor embodied navigation. NavVerse contains 100 indoor scenes, 50 urban outdoor scenes, and 50 indoor-to-outdoor scenes, and 10,000 episodes spanning Object Navigation, Vision-and-Language Navigation, and Place Navigation tasks, where agents search for semantic points of interest such as restaurants or banks. Agents are evaluated through executable robot interfaces using task-success, path-efficiency, and safety metrics. Zero-shot experiments with RL, VLA, and modular baselines show that current agents remain far from solving cross-context navigation: end-to-end VLAs obtain the highest zero-shot success, while the modular method provides the strongest safety profile. PlaceNav further reveals a clear drop from outdoor to indoor-to-outdoor scenes, indicating that adaptation remains major bottleneck.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。