让机器人记住看不见的物体位置,实现更智能的远程操作。
Spatial Memory for Out-of-Vision Manipulation in Vision-Language-Action

- 用多视角扫描构建持久空间记忆,突破视觉范围限制。
- 在5个真实场景中提升任务成功率,实现近一次抓取即定位目标。
- 适合需要长时间记忆与跨视野操作的机器人研究者使用。
我们提出SOMA——一种面向视觉-语言-动作(VLA)模型的出视野操作空间记忆框架。现有VLA模型通常假设目标始终可见,导致目标脱离视野时行为僵化。SOMA通过可移动摄像头获取多视角观测,构建持续的空间记忆,使模型能超越当前视觉视野进行推理。该框架包含三部分:空间记忆构建,通过扫描将角度信息融合为统一的时空语义表示;动态记忆精炼,维持长期全局一致性;上下文记忆检索,在操作中激活与指令相关的空间线索。我们在五个具有挑战性的真实世界出视野操作任务上评估SOMA,包括多步和双臂场景,目标初始不可见。实验表明,SOMA不仅显著提升任务成功率,还带来质的改变:目标定位更快、视角搜索减少、在部分可观测条件下实现近一击抓取。在RoboCasa GR1和SimplerEnv上的附加实验进一步验证了其记忆设计在常规全可观测设置下的有效性。代码即将开源。
原文摘要 · Abstract (English)
We introduce SOMA, the Spatial Memory framework for Out-of-Vision Manipulation in Vision-Language-Action (VLA) models. Most existing VLAs implicitly assume that task-relevant objects are always visible, leading to brittle and reactive behaviors when targets fall outside the camera's field of view. SOMA addresses this limitation by equipping VLAs with a persistent spatial memory constructed from multi-view observations acquired via a movable head camera, enabling reasoning beyond the current visual frustum. The framework consists of three components: Spatial Memory Construction, which aggregates angular-wise observations into a unified spatial-semantic representation through scanning; Dynamic Memory Refinement, which maintains global consistency over time; and Contextual Memory Retrieval, which activates instruction-relevant spatial cues during manipulation. We evaluate SOMA on five challenging real-world out-of-vision manipulation tasks, including multi-step and dual-arm scenarios where target objects are initially invisible. Experimental results show that SOMA not only improves task success rates, but also induces qualitatively different manipulation behaviors, with faster target localization, reduced viewpoint search, and near one-shot grasping under partial observability. Additional experiments on RoboCasa GR1 and SimplerEnv further validate the effectiveness of SOMA's memory design under conventional fully observable settings. Code will be released soon.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。