仅用单目相机实现零样本语义导航,能快速适应新环境并高效找物。
RANGER: A Monocular Zero-Shot Semantic Navigation Framework through Visual Contextual Adaptation
- 通过短时视频感知环境,无需预建图或深度信息
- 在HM3D上导航成功率超基线,探索效率提升显著
- 适合真实场景中无先验地图的机器人自主寻物任务
复杂环境中高效定位目标与自主导航是具身应用的核心挑战。尽管多模态基础模型已支持零样本物体目标导航(无需微调),现有方法仍存在两大局限:(1) 严重依赖真值深度和位姿信息,限制真实场景应用;(2) 缺乏视觉上下文学习能力,无法从短时遍历视频中提取几何与语义先验。为此,我们提出RANGER——一种仅依赖单目相机的零样本、开放词汇语义导航框架。借助强大3D基础模型,RANGER消除对深度与位姿的依赖,并具备强视觉上下文学习能力。仅需观察一段目标环境短视频,系统即可显著提升任务效率,且无需架构修改或任务特定重训练。框架集成关键帧3D重建、语义点云生成、视觉语言模型驱动的探索价值估计、高层自适应路径点选择与低层动作执行等模块。在HM3D基准与真实环境中的实验表明,RANGER在导航成功率与探索效率方面表现优异,且展现出卓越的上下文适应能力,无需此前3D地图。
原文摘要 · Abstract (English)
Efficient target localization and autonomous navigation in complex environments are fundamental to real-world embodied applications. While recent advances in multimodal foundation models have enabled zero-shot object goal navigation, allowing robots to search for arbitrary objects without fine-tuning, existing methods face two key limitations: (1) heavy reliance on ground-truth depth and pose information, which restricts applicability in real-world scenarios; and (2) lack of visual in-context learning (VICL) capability to extract geometric and semantic priors from environmental context, as in a short traversal video. To address these challenges, we propose RANGER, a novel zero-shot, open-vocabulary semantic navigation framework that operates using only a monocular camera. Leveraging powerful 3D foundation models, RANGER eliminates the dependency on depth and pose while exhibiting strong VICL capability. By simply observing a short video of the target environment, the system can also significantly improve task efficiency without requiring architectural modifications or task-specific retraining. The framework integrates several key components: keyframe-based 3D reconstruction, semantic point cloud generation, vision-language model (VLM)-driven exploration value estimation, high-level adaptive waypoint selection, and low-level action execution. Experiments on the HM3D benchmark and real-world environments demonstrate that RANGER achieves competitive performance in terms of navigation success rate and exploration efficiency, while showing superior VICL adaptability, with no previous 3D mapping of the environment required.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。