首个真人验证的3D室内导航基准,支持四级语义目标定位。
LangMap: A Human-Verified Benchmark for Hierarchical Open-Vocabulary Goal Navigation
- 构建四层级语义目标体系,含场景、房间、区域和实例
- 18000+任务,实例描述匹配率比GOAT-Bench高23个百分点
- 仅用图像实现顶尖导航性能,适合多目标与无深度场景研究
语言条件下的目标导航(LGN)要求智能体在无步骤引导下定位用户指定目标。现有基准多聚焦类别级目标或依赖视觉-语言模型生成的实例描述,常含歧义和语义错误,影响评估可靠性。我们提出HieraNav,一种开放词汇的分层语义目标导航任务,目标涵盖场景、房间、区域和实例四个层级。为此,我们构建了LangMap——首个基于真实3D室内场景、经真人验证语义标注的导航基准,支持全层级任务。该数据集包含414个物体类别,通过对比同场景区域与实例的严格对照标注协议生成区域标签及可区分的区域与实例描述,共覆盖超过18,000个任务。每个目标配备简洁与详细双版本描述,支持多风格指令评估。定量与定性分析验证了标注质量;尤其,实例描述在文本到视角匹配上较GOAT-Bench提升23个百分点。我们进一步提出PlaNaVid,一个仅使用RGB的强基线模型,结合有界多样性记忆(BDM)与高层规划,驱动反应式策略完成多目标导航,在无深度、3D场景图或物体掩码条件下达到顶尖成功率。分析表明,记忆机制与丰富上下文显著提升性能,而长尾类别、小物体、远距离目标及多目标完成仍是挑战。基准已开源:https://bo-miao.github.io/LangMap
原文摘要 · Abstract (English)
Language-conditioned goal navigation (LGN) requires agents to locate user-specified targets without step-by-step guidance. However, existing benchmarks largely focus on category-level goals or rely on instance descriptions generated by vision-language models (VLMs), which often contain ambiguities and semantic errors, limiting systematic and reliable evaluation. We introduce HieraNav, an open-vocabulary LGN task with goals specified at four hierarchical semantic levels: scene, room, region, and instance. To this end, we present Language as a Map (LangMap), to our knowledge the first real-world 3D indoor navigation benchmark with human-verified semantic annotations to support tasks across all four goal levels. LangMap provides region labels and discriminative region and instance descriptions covering 414 object categories, produced through a rigorous contrastive annotation protocol comparing same-scene regions and instances, and contains over 18K tasks. Each target is paired with concise and detailed descriptions, enabling evaluation across instruction styles. Quantitative and qualitative analyses validate our annotation quality; notably, our instance descriptions outperform GOAT-Bench annotations by 23 percentage points in text-to-view matching. We further introduce PlaNaVid, a strong RGB-only baseline that combines Bounded Diverse Memory (BDM) with high-level planning to prime a reactive policy for multi-goal navigation. PlaNaVid achieves top-tier success rates without depth, 3D scene representations, or object masks. Further analysis shows that memory and richer context boost performance, while long-tailed categories, small objects, distant targets, and multi-goal completion remain open challenges. The benchmark is available at https://bo-miao.github.io/LangMap
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。