让机器人记住长期导航经历并回答时空问题
ReMEmbR: Building and Reasoning Over Long-Horizon Spatio-Temporal Memory for Robot Navigation
- 用检索增强记忆构建时空结构化历史
- 在长时序视频问答中超越大模型基线
- 适合需要长期记忆的机器人交互场景
在长时间跨度内导航与理解复杂环境是机器人的重大挑战。人类用户可能询问某事件发生的位置、时间或距今多久,这要求机器人能对部署历史进行长周期推理。为此,我们提出面向具身机器人的检索增强记忆系统 ReMEmbR,用于机器人导航中的长时序视频问答。为评估 ReMEmbR,我们构建了 NaVQA 数据集,对长时序机器人导航视频标注了空间、时间和描述性问题。ReMEmbR 采用分阶段结构:记忆构建与查询阶段,融合时间、空间信息和图像,高效处理持续增长的机器人历史。实验表明,ReMEmbR 在性能上优于大语言模型(LLM)和视觉语言模型(VLM)基线,实现低延迟的长时序推理。此外,我们在真实机器人上部署验证,证明该方法可应对多样查询。相关数据集、代码、视频等资源详见:https://nvidia-ai-iot.github.io/remembr
原文摘要 · Abstract (English)
Navigating and understanding complex environments over extended periods of time is a significant challenge for robots. People interacting with the robot may want to ask questions like where something happened, when it occurred, or how long ago it took place, which would require the robot to reason over a long history of their deployment. To address this problem, we introduce a Retrieval-augmented Memory for Embodied Robots, or ReMEmbR, a system designed for long-horizon video question answering for robot navigation. To evaluate ReMEmbR, we introduce the NaVQA dataset where we annotate spatial, temporal, and descriptive questions to long-horizon robot navigation videos. ReMEmbR employs a structured approach involving a memory building and a querying phase, leveraging temporal information, spatial information, and images to efficiently handle continuously growing robot histories. Our experiments demonstrate that ReMEmbR outperforms LLM and VLM baselines, allowing ReMEmbR to achieve effective long-horizon reasoning with low latency. Additionally, we deploy ReMEmbR on a robot and show that our approach can handle diverse queries. The dataset, code, videos, and other material can be found at the following link: https://nvidia-ai-iot.github.io/remembr
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。