让机器人自由选路,靠语言和视觉动态导航
DyNaVLM: Zero-Shot Vision-Language Navigation System with Dynamic Viewpoints and Self-Refining Graph Memory
- 用动态视角和语言推理决定走哪条路
- 记忆图可自我更新,支持多机器人共享
- 无需训练就能在真实世界中稳定运行
我们提出DyNaVLM,一种基于视觉语言模型(VLM)的端到端视觉语言导航框架。与以往受限于固定角度或距离间隔的方法不同,本系统使智能体可通过视觉-语言推理自由选择导航目标。核心是自优化图记忆结构,能:1)以可执行的拓扑关系存储物体位置;2)通过分布式图更新实现跨机器人记忆共享;3)通过检索增强提升VLM决策能力。系统无需任务特定训练或微调,在GOAT和ObjectNav基准上表现优异。真实世界测试进一步验证其鲁棒性与泛化能力。三大创新——动态动作空间设计、协同图记忆机制、零训练部署——为可扩展具身机器人建立新范式,弥合离散视觉语言导航任务与连续现实导航之间的差距。
原文摘要 · Abstract (English)
We present DyNaVLM, an end-to-end vision-language navigation framework using Vision-Language Models (VLM). In contrast to prior methods constrained by fixed angular or distance intervals, our system empowers agents to freely select navigation targets via visual-language reasoning. At its core lies a self-refining graph memory that 1) stores object locations as executable topological relations, 2) enables cross-robot memory sharing through distributed graph updates, and 3) enhances VLM's decision-making via retrieval augmentation. Operating without task-specific training or fine-tuning, DyNaVLM demonstrates high performance on GOAT and ObjectNav benchmarks. Real-world tests further validate its robustness and generalization. The system's three innovations: dynamic action space formulation, collaborative graph memory, and training-free deployment, establish a new paradigm for scalable embodied robot, bridging the gap between discrete VLN tasks and continuous real-world navigation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。