arXiv:2511.11478cs.ROcs.CV2025-11中稿 · AAAI被引 24

提出新框架让机器人记住物体历史,解决复杂操作中的记忆难题。

Rethinking Progression of Memory State in Robotic Manipulation: An Object-Centric Perspective

  • 用空间槽机制保持物体身份一致,支持长期跟踪。
  • 在数百帧任务中表现优于传统模型,避免重复动作。
  • 适合需要物体历史推理的机器人控制场景。

随着具身智能体在日益复杂的环境中运行,对个体物体实例随时间的感知、跟踪和推理能力变得至关重要,尤其在涉及视觉相似物体的序列交互任务中。在非马尔可夫环境下,关键决策线索常隐藏于物体特定的历史记录中,而非当前场景。若缺乏对先前交互(何时、何地、如何变化)的持久记忆,视觉-运动策略可能失效、重复动作或忽略已完成步骤。为此,我们引入LIBERO-Mem,一个用于在物体级部分可观测条件下压力测试机器人操作的非马尔可夫任务套件,结合短时与长时物体追踪及时间序列子目标,要求超越当前帧的推理。然而,视觉-语言-动作(VLA)模型在此类设置中往往表现不佳,即使仅数百帧的任务也面临令牌扩展不可行的问题。我们提出Embodied-SlotSSM,一种面向时间可扩展性的槽中心VLA框架。它通过两个机制维持时空一致的槽身份:(1) 槽状态空间建模以重构短期历史;(2) 关系编码器将输入标记与动作解码对齐。二者协同实现时序锚定、上下文感知的动作预测。实验表明,Embodied-SlotSSM在LIBERO-Mem及通用任务上均达基线性能,为物体中心的非马尔可夫推理提供可扩展解决方案。

原文摘要 · Abstract (English)

As embodied agents operate in increasingly complex environments, the ability to perceive, track, and reason about individual object instances over time becomes essential, especially in tasks requiring sequenced interactions with visually similar objects. In these non-Markovian settings, key decision cues are often hidden in object-specific histories rather than the current scene. Without persistent memory of prior interactions (what has been interacted with, where it has been, or how it has changed) visuomotor policies may fail, repeat past actions, or overlook completed ones. To surface this challenge, we introduce LIBERO-Mem, a non-Markovian task suite for stress-testing robotic manipulation under object-level partial observability. It combines short- and long-horizon object tracking with temporally sequenced subgoals, requiring reasoning beyond the current frame. However, vision-language-action (VLA) models often struggle in such settings, with token scaling quickly becoming intractable even for tasks spanning just a few hundred frames. We propose Embodied-SlotSSM, a slot-centric VLA framework built for temporal scalability. It maintains spatio-temporally consistent slot identities and leverages them through two mechanisms: (1) slot-state-space modeling for reconstructing short-term history, and (2) a relational encoder to align the input tokens with action decoding. Together, these components enable temporally grounded, context-aware action prediction. Experiments show Embodied-SlotSSM's baseline performance on LIBERO-Mem and general tasks, offering a scalable solution for non-Markovian reasoning in object-centric robotic policies.

机器人操作记忆机制物体中心

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。