让机器人长期记忆更智能,能精准回答复杂任务指令。
STaR: Scalable Task-Conditioned Retrieval for Long-Horizon Multimodal Robot Memory
- 构建多模态长时记忆,支持细粒度环境语义保留。
- 提出可扩展的任务条件检索算法,降低冗余提升信息密度。
- 在真实机器人上验证,适用于室内外复杂场景。
移动机器人常在多样、动态的开放环境中长期运行,如仓库、制造车间、农田和道路。核心挑战在于建立可扩展的长时记忆系统,以支持规划、检索和推理等代理式工作流,针对不同粒度的开放式指令生成精确、可执行的导航答案。本文提出STaR,一种代理式推理框架:(i) 构建与任务无关的多模态长期记忆,能泛化到未见查询,同时保留物体属性、空间关系和动态事件等细粒度环境语义;(ii) 基于信息瓶颈原理,提出可扩展的任务条件检索算法,从长期记忆中提取紧凑、非冗余、信息丰富的候选记忆集合,用于上下文推理。我们在NaVQA(混合室内外校园场景)和WH-VQA(基于Isaac Sim构建的定制化仓库基准,含大量视觉相似物体)上评估,两个数据集上均显著优于强基线,成功率更高,空间误差大幅降低。进一步在真实Husky轮式机器人上部署,验证了其在室内外环境中的鲁棒长时推理能力、可扩展性与实用性。项目官网:https://trailab.github.io/STaR-website/
原文摘要 · Abstract (English)
Mobile robots are often deployed over long durations in diverse open, dynamic scenes, including indoor setting such as warehouses and manufacturing facilities, and outdoor settings such as agricultural and roadway operations. A core challenge is to build a scalable long-horizon memory that supports an agentic workflow for planning, retrieval, and reasoning over open-ended instructions at variable granularity, while producing precise, actionable answers for navigation. We present STaR, an agentic reasoning framework that (i) constructs a task-agnostic, multimodal long-term memory that generalizes to unseen queries while preserving fine-grained environmental semantics (object attributes, spatial relations, and dynamic events), and (ii) introduces a Scalable Task Conditioned Retrieval algorithm based on the Information Bottleneck principle to extract from long-term memory a compact, non-redundant, information-rich set of candidate memories for contextual reasoning. We evaluate STaR on NaVQA (mixed indoor/outdoor campus scenes) and WH-VQA, a customized warehouse benchmark with many visually similar objects built with Isaac Sim, emphasizing contextual reasoning. Across the two datasets, STaR consistently outperforms strong baselines, achieving higher success rates and markedly lower spatial error. We further deploy STaR on a real Husky wheeled robot in both indoor and outdoor environments, demonstrating robust long horizon reasoning, scalability, and practical utility. Project Website: https://trailab.github.io/STaR-website/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。