用自演化记忆让机器人长期自主导航,接近100%成功率。
AllDayNav: Lifelong Navigation via Real-World Reinforcement Learning

- 通过强化学习将场景动态隐式编码到大模型参数中
- 跨房间、跨任务场景下成功率逼近100%,路径更高效
- 适合需要长期适应真实动态环境的机器人导航研究
在动态环境中实现持续性的具身导航,要求机器人从零散观测中形成持久的场景理解,现有依赖显式地图或场景图的方法难以泛化至非结构化场景。我们提出AllDayNav,一种基于强化学习的终身自学习导航框架,通过自演化多模态记忆,隐式将场景动态编码进千亿级参数的大模型中,该记忆可自主维护并更新视觉关键帧、语义描述和时间上下文,同时生成开放词汇指令、图像目标与结构化奖励。在合成与真实世界环境中进行的跨房间、跨回合、跨任务实验表明,AllDayNav的成功率接近100%,在路径效率与鲁棒性上持续超越强基准方法,证明了隐式、记忆驱动的强化学习是可靠终身导航的可扩展替代方案。
原文摘要 · Abstract (English)
Lifelong embodied navigation in dynamic environments requires robots to form persistent scene understanding from fragmentary observations, which remains difficult for existing methods that rely on explicit maps or scene graphs and struggle to generalize beyond structured settings. We propose AllDayNav, a lifelong self-learning navigation framework that implicitly encodes scene dynamics into the billion-scale parameters of a large model via reinforcement learning, powered by a self-evolving multimodal memory that maintains and updates visual keyframes, semantic descriptions, and temporal context while autonomously generating open-vocabulary instructions, image goals, and structured rewards. Experiments in both synthetic and real-world environments across cross-room, cross-episode, and cross-task scenarios show that AllDayNav achieves success rates approaching $100\%$ and consistently surpasses strong map-based, VLM, and RL baselines in path efficiency and robustness, demonstrating implicit, memory-driven reinforcement learning as a scalable alternative to explicit mapping for reliable lifelong navigation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。