让机器人听懂指令后主动看、记得住,导航更准
MAG-Nav: Language-Driven Object Navigation Leveraging Memory-Reserved Active Grounding
- 用动态视角和记忆回溯提升视觉理解能力
- 零样本下在未知环境导航成功率显著提升
- 适合需要强泛化能力的智能机器人场景
仅依赖自然语言描述进行未知环境中的视觉导航是智能机器人的关键能力。本文提出一种基于现成视觉语言模型(VLMs)的导航框架,引入两种类人机制:基于视角的主动定位,可动态调整机器人视角以优化视觉观察;历史记忆回溯,使系统能保留并重新评估不确定的感知信息。与被动依赖偶然视觉输入的方法不同,本方法主动优化感知,并利用记忆化解歧义,在复杂未见环境中显著提升视觉-语言对齐效果。该框架无需标注数据或模型微调,实现零样本泛化,可应对多样化、开放式的语言指令。在Habitat-Matterport 3D(HM3D)上的实验表明,其性能优于现有最优方法。此外,通过四足机器人的真实部署验证了其实用性,实现了鲁棒有效的导航表现。
原文摘要 · Abstract (English)
Visual navigation in unknown environments based solely on natural language descriptions is a key capability for intelligent robots. In this work, we propose a navigation framework built upon off-the-shelf Visual Language Models (VLMs), enhanced with two human-inspired mechanisms: perspective-based active grounding, which dynamically adjusts the robot's viewpoint for improved visual inspection, and historical memory backtracking, which enables the system to retain and re-evaluate uncertain observations over time. Unlike existing approaches that passively rely on incidental visual inputs, our method actively optimizes perception and leverages memory to resolve ambiguity, significantly improving vision-language grounding in complex, unseen environments. Our framework operates in a zero-shot manner, achieving strong generalization to diverse and open-ended language descriptions without requiring labeled data or model fine-tuning. Experimental results on Habitat-Matterport 3D (HM3D) show that our method outperforms state-of-the-art approaches in language-driven object navigation. We further demonstrate its practicality through real-world deployment on a quadruped robot, achieving robust and effective navigation performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。