用视觉嵌入+时空定位,让机器人长期记忆更准更快
RAVEN: Long-Horizon Reasoning & Navigation with a Visuo-Spatio-Temporal Memory

- 直接存视觉特征+位姿时间,避免图像转文字的损失
- 在多个场景中检索成本低10倍,效果媲美顶级大模型
- 已在真实机器人上实现跨大空间的自然语言导航
长期部署的机器人需要一种紧凑且可扩展的记忆系统,以保留精细的视觉语义,将观测结果锚定在时空之中,并支持高效存储与检索。本文提出RAVEN,一种用于长时序机器人问答与导航的智能体记忆系统。RAVEN将视觉嵌入与位姿、时间信息一同存入向量数据库,并通过空间地图实现检索定位,以回答问题和导航至目标。通过直接操作视觉嵌入,RAVEN避免了有损的图像到文本描述转换,实现了大规模下的精准语义、空间与时间检索。在多个模拟与真实世界的视频问答基准上,RAVEN始终优于基于描述的记忆系统,在长时序任务上达到前沿视觉语言模型的表现,但检索成本降低10倍。最后,我们在Unitree Go1机器人上实现了RAVEN,成功完成多个大型室内环境中的长时序自然语言目标到达任务。
原文摘要 · Abstract (English)
Long-term robot deployment requires a compact and scalable memory that preserves fine-grained visual semantics, grounds observations in space and time, and enables efficient storage and retrieval. In this paper, we propose RAVEN, an agentic memory system for long-horizon robotic question answering and navigation. RAVEN stores visual embeddings with pose and time in a vector database, and grounds retrieval in a spatial map to answer queries and navigate to goals. By operating directly on visual embeddings, RAVEN avoids lossy image-to-text captioning and enables accurate semantic, spatial, and temporal retrieval at scale. Across several simulated and real-world video question-answering benchmarks, RAVEN consistently surpasses caption-based memory systems and matches frontier VLMs on long-horizon tasks at 10$\times$ lower retrieval cost. Finally, we instantiate RAVEN on a Unitree Go1 robot for the task of long-horizon navigation for natural language goal-reaching, and show successful deployment over several large indoor environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。