让机器人像人一样记过去、想未来,提升长时序操作能力。
MemoryVLA++: Temporal Modeling via Memory and Imagination in Vision-Language-Action Models

- 用记忆库和世界模型模拟人类工作记忆与预判能力
- 真实机器人任务中性能提升最高达28%
- 适合需要长期规划与状态预测的机器人研究者
时间建模对机器人操作至关重要,有效控制需依赖对过往交互的记忆与对未来状态的想象。然而,多数视觉-语言-动作(VLA)模型仅基于当前观测,难以应对长时序、时序依赖任务。认知科学表明,人类依赖工作记忆缓冲短期上下文、海马系统保存事件记忆,并通过内在模型预判未来状态演变。受此启发,我们提出MemoryVLA++,一种完整的时序建模框架,为VLA模型赋予记忆与想象能力。预训练视觉-语言模型将当前观测编码为感知与认知令牌,形成工作记忆;这些令牌查询感知-认知记忆库,以获取相关历史上下文。该记忆库存储过往交互中的低层细节与高层语义,通过冗余感知整合机制更新。一个世界模型在去噪潜空间中想象未来状态,想象出的潜变量在记忆引导下融合,生成全时序感知令牌。最终令牌用于条件化扩散动作专家,预测时序一致的动作序列。我们在5个仿真基准与3类真实机器人任务(共3台机器人)上进行广泛实验,涵盖通用操作、长时序任务、鲁棒性与泛化能力。结果表明,该方法在Libero、SimplerEnv、Mikasa-Robo、Calvin、Libero-Plus及多样化真实任务中表现优异,在真实机器人上,通用任务、记忆依赖任务与想象依赖任务分别提升+9%、+26%、+28%。
原文摘要 · Abstract (English)
Temporal modeling is essential for robotic manipulation, as effective control requires both memory of past interactions and imagination of future states. However, most VLA models rely primarily on the current observation and therefore struggle with long-horizon, temporally dependent tasks. Cognitive science suggests that humans rely on working memory to buffer short-lived context, the hippocampal system to preserve episodic memory of past experience, and internal models to imagine possible future state evolution. Inspired by these mechanisms, we propose MemoryVLA++, a full temporal modeling framework that equips VLA models with memory and imagination for robotic manipulation. A pretrained VLM encodes the current observation into perceptual and cognitive tokens, forming working memory. These tokens query a Perceptual-Cognitive Memory Bank to retrieve relevant historical context. This bank stores low-level details and high-level semantics from past interactions, and is updated through redundancy-aware consolidation. A world model imagines future states in a denoising latent space, and the imagined latents are integrated under memory guidance to form full temporal-aware tokens. The resulting tokens condition a diffusion action expert to predict temporally consistent action sequences. We conduct extensive experiments on 5 simulation benchmarks and 3 categories of real-robot tasks across 3 robots, covering general manipulation, long-horizon temporal tasks, robustness, and generalization. Our method achieves strong performance across Libero, SimplerEnv, Mikasa-Robo, Calvin, Libero-Plus, and diverse real-robot tasks, validating the effectiveness of full temporal modeling with memory and imagination. For example, on real robots, it achieves +9%, +26%, +28% gains on general, memory-dependent, and imagination-dependent tasks. Project Page: https://shihao1895.github.io/MemoryVLA-PP-Web
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。