arXiv:2608.10886cs.CVcs.RO2026-08

让机器人理解人如何在动态环境中使用物品并组合成目标行为。

GESTO: Human-Centric Spatio-Temporal Memory for Reasoning in Dynamic Scenes

论文配图:GESTO: Human-Centric Spatio-Temporal Memory for Reasoning in Dynamic Scenes
图 1 · 摘自论文原文
  • 构建分层时空记忆,融合物体位置与人物交互事件
  • 在标准任务上达0.71~0.75分数,接近有真值输入的模型
  • 适合需要理解人类活动逻辑的机器人场景推理

在人类环境运行的机器人需要记忆不仅包含物体及其位置,还需记录人们随时间对物品的使用方式,以及个体互动如何构成目标导向的行为。现有4D场景图保留了物体和空间的历史,但忽略了行为结构;而行为表示要么未与持续的3D场景关联,要么依赖外部提供的事件边界和物体关联。我们提出GESTO(Grounded Event and Spatio-Temporal memOry),一种将持久4D场景图与原子级人-物交互及目标驱动事件的双层结构耦合的时空记忆。从RGB-D观测流中,GESTO自动提取带时间戳的交互,将其锚定到持久场景实体,聚类为事件,并利用事件上下文优化不确定的物体关联。一个关系感知的工具调用代理查询该记忆以实现以行为为中心的时空推理。我们在现有基准的可复现文本、二分类和时间类别上评估GESTO,另增加40个Space2Event和Event2Space查询。GESTO在标准类别上得分0.71、0.75、0.70,接近提供真值事件和物体锚定的方法,且显著优于移除这些输入的相同推理框架。其在Space2Event和Event2Space查询上分别达0.73和0.75。消融实验表明,层次化事件结构与上下文感知的锚定精炼提供互补优势,支持以行为为基础的层次化记忆用于动态人类环境中的回溯推理。

原文摘要 · Abstract (English)

Robots operating in human environments need memories that capture not only what objects exist and where, but also how people use them over time and how individual interactions compose into goal-directed activities. Existing 4D scene graphs preserve object and place histories but omit activity structure, whereas activity representations are either not grounded in persistent 3D scenes or rely on externally provided event boundaries and object associations. We present GESTO (Grounded Event and Spatio-Temporal memOry), a spatio-temporal memory that couples a persistent 4D scene graph with a two-level hierarchy of atomic human--object interactions and goal-driven events. From an RGB-D observation stream, GESTO automatically extracts timestamped interactions, grounds them to persistent scene entities, groups them into events, and uses event context to refine uncertain object associations. A relation-aware tool-calling agent queries the resulting memory for activity-centric spatio-temporal reasoning. We evaluate GESTO on the reproducible text, binary, and time categories of an existing benchmark, together with 40 new Space2Event and Event2Space queries. GESTO achieves scores of 0.71, 0.75, and 0.70 on the standard categories, approaching a method supplied with ground-truth event and object grounding, while substantially outperforming the same reasoning framework when these inputs are removed. It further achieves 0.73 and 0.75 on Space2Event and Event2Space queries. Ablations show that hierarchical event structure and context-aware grounding refinement provide complementary benefits, supporting activity-grounded hierarchical memory for retrospective reasoning in dynamic human environments.

机器人时空记忆行为理解人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。