arXiv:2511.14004cs.RO2025-11被引 3

统一时空搜索,让机器人能记住过去、理解位置、自主行动找物品。

Searching in Space and Time: Unified Memory-Action Loops for Open-World Object Retrieval

  • 用记忆与视觉语言模型联动,动态决定是查时间还是查空间。
  • 在模拟和真实机器人上表现优于传统方法,尤其在复杂场景中提升显著。
  • 适合需要长期记忆与环境交互的智能机器人应用。

服务机器人需在动态开放世界中检索物品,请求可能涉及属性(如‘红色杯子’)、空间上下文(如‘桌子上的杯子’)或过去状态(如‘昨天在这里的杯子’)。现有方法仅覆盖部分问题:场景图捕捉空间关系但忽略时间定位,时序推理模型处理动态变化但不支持具身交互,动态场景图虽兼顾二者却局限于封闭词汇集。我们提出STAR(SpatioTemporal Active Retrieval)框架,将记忆查询与具身动作整合于单一决策循环。STAR利用非参数化长期记忆与工作记忆实现高效回溯,并通过视觉-语言模型在每一步选择时间或空间动作。我们构建STARBench基准,涵盖模拟与真实环境中的时空物体搜索任务。在STARBench及Tiago机器人上的实验表明,STAR持续优于场景图与纯记忆基线,证明将时空搜索视为统一问题的价值。

原文摘要 · Abstract (English)

Service robots must retrieve objects in dynamic, open-world settings where requests may reference attributes ("the red mug"), spatial context ("the mug on the table"), or past states ("the mug that was here yesterday"). Existing approaches capture only parts of this problem: scene graphs capture spatial relations but ignore temporal grounding, temporal reasoning methods model dynamics but do not support embodied interaction, and dynamic scene graphs handle both but remain closed-world with fixed vocabularies. We present STAR (SpatioTemporal Active Retrieval), a framework that unifies memory queries and embodied actions within a single decision loop. STAR leverages non-parametric long-term memory and a working memory to support efficient recall, and uses a vision-language model to select either temporal or spatial actions at each step. We introduce STARBench, a benchmark of spatiotemporal object search tasks across simulated and real environments. Experiments in STARBench and on a Tiago robot show that STAR consistently outperforms scene-graph and memory-only baselines, demonstrating the benefits of treating search in time and search in space as a unified problem. For more information: https://amrl.cs.utexas.edu/STAR.

机器人时空搜索具身智能记忆机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。