提出分层记忆门控模型,提升机器人长任务操作的鲁棒性与记忆能力。
HiMem-WAM: Hierarchical Memory-Gated World Action Models for Robotic Manipulation

- 分层潜变量联合学习动作与技能,实现时序结构抽象。
- 在边界处触发记忆更新,减少冗余信息并支持因果推理。
- 在真实场景和多个基准上表现优异,尤其适合长时序操作任务。
世界动作模型(WAMs)作为具身智能的新范式,通过学习与动作相关的视觉动态,显著提升了泛化能力和鲁棒性。然而,现有WAMs在长时序机器人操作中仍面临任务相关记忆不足的问题。为此,本文提出HiMem-WAM,一种分层记忆门控的WAM,整合了以运动为中心的潜动作、高层技能潜变量以及边界触发的记忆更新机制。具体而言,构建分层潜动作框架,联合学习低层运动与高层技能潜变量,实现结构化的时序抽象;同时设计边界感知记忆门,仅在预测的技能转换点写入紧凑的任务状态,避免测试时生成未来视频或光流估计,支持因果推理。在LIBERO、LIBERO-PLUS、RMBench及真实世界任务上的评估表明,分层潜变量增强了部署扰动下的鲁棒性,记忆模块显著提升了依赖记忆的长时序操作性能。
原文摘要 · Abstract (English)
World Action Models (WAMs) have emerged as a new powerful paradigm for embodied intelligence, learning action-relevant visual dynamics that significantly enhance generalization and robustness. However, existing WAMs still struggle with task-relevant memory in long-horizon robotic manipulation. To address this, we present HiMem-WAM, a Hierarchical Memory-Gated WAM that integrates motion-centric latent actions, high-level skill latents, and boundary-triggered memory updates. Specifically, we develop a hierarchical latent action framework that jointly learns low-level motion and high-level skill latents, providing structured temporal abstraction. Meanwhile, a boundary-aware memory gate writes compact task states at predicted skill transitions, enabling causal inference without test-time generation of future video or optical flow estimation. Evaluated on LIBERO, LIBERO-PLUS, RMBench and real-world tasks, HiMem-WAM shows that hierarchical latents improve robustness under deployment perturbations, and the memory module substantially benefits memory-dependent long-horizon manipulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。