arXiv:2607.07608cs.ROcs.CV2026-07被引 1

让机器人记住过去动作,更智能地完成复杂任务。

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation

论文配图:Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation
图 1 · 摘自论文原文
  • 把历史经验变成可直接参与推理的潜空间记忆
  • 在模拟和真实机器人任务中显著提升长序列操作能力
  • 适合需要连续决策的机器人控制研究者

主流视觉-语言-动作(VLA)模型在马尔可夫假设下仅依赖当前观测预测动作,难以应对长时间、依赖时间的任务。现有带记忆的VLA模型或扩大观察窗口,或从记忆库检索信息作为辅助上下文,但记忆未融入VLA本征潜空间,导致历史经验无法与多模态推理和动作生成自然融合。为此,我们提出LaMem-VLA,一种原生潜空间记忆框架,将历史经验重构为潜记忆标记,并直接嵌入VLA推理流程。其核心包含四个协同组件:(i) 管理员将历史经验分为短期和长期记忆存储;(ii) 检索器利用多模态认知查询两个存储库;(iii) 压缩器将检索到的信息压缩为紧凑的短时和长时潜记忆标记;(iv) 编织器将这些记忆标记与当前观测和指令整合成连续嵌入序列。通过在统一连续潜空间中表示、检索和消费历史经验,LaMem-VLA实现记忆对推理与动作生成的直接引导,在SimplerEnv和LIBERO上的实验验证了其优越性。

原文摘要 · Abstract (English)

Mainstream Vision-Language-Action (VLA) models predict actions primarily from the current observation under a Markovian assumption, thus struggling with long-horizon, temporally dependent tasks. Existing memory-augmented VLAs either expand the observation window or retrieve history from the memory bank as auxiliary policy-side context. However, they leave memory outside the native latent embedding space of VLA reasoning, preventing historical experience from being fluidly interleaved with multimodal reasoning and action formation. To this end, we introduce LaMem-VLA, a latent-memory-native framework that reconstructs historical experience into latent memory tokens and directly interweaves them with VLA reasoning. At its core, LaMem-VLA introduces four coordinated components: (i) a curator that organizes historical experience into two complementary short-term and long-term memory vaults; (ii) a seeker that queries both vaults using the multimodal cognition to retrieve context-relevant evidence; (iii) a condenser that reconstructs the retrieved evidence into compact short-term and long-term latent memory tokens; and (iv) a weaver that injects these memory tokens with the current observation and instruction into one continuous embedding sequence. By representing, retrieving, and consuming historical experience entirely in the same continuous latent space, LaMem-VLA enables memory to directly participate in VLA reasoning and guide action generation under a bounded context. Extensive experiments on SimplerEnv and LIBERO demonstrate the superiority of our LaMem-VLA.

机器人控制记忆机制多模态推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。