让机器人模型记住过去视角,实现长时间精准操作预测。
Mem-World: Memory-Augmented Action-Conditioned World Models for Persistent Robot Manipulation

- 用4D表面点索引记忆库,动态追踪物体在时间与空间中的变化。
- 在复杂操作中生成更持久的视频预测,真实性能相关性提升14.5%。
- 适合需要长期记忆和高精度视觉预测的机器人任务研究者。
动作条件世界模型为机器人学习提供了可扩展的替代方案,通过生成一致动作的视频序列减少真实实验成本。然而,在持续操作任务中,频繁的末端执行器遮挡和快速腕部摄像头运动导致当前观测不足以预测未来视图,使模型容易遗忘或幻化早期帧中的场景细节。现有记忆检索策略在动态操作场景中常无法识别有效历史信息。为此,我们提出Mem-World——一种增强记忆的多视角动作条件世界模型。核心是W-VMem,一种以腕部视角为中心、基于表面点(surfel)索引的4D记忆结构,将历史观测锚定在随时间演化的表面元素上。通过显式建模场景元素被观测的时间与位置,W-VMem实现了基于未来动作的几何感知历史检索。生成时,通过基于表面点的渲染与评分选择相关历史帧,提供有信息量且无冗余的上下文用于预测。大量实验表明,Mem-World能在复杂操作场景中生成持久的视频滚动,相比Ctrl-World提升政策评估可靠性,与真实世界表现的皮尔逊相关性提高14.5%,并通过合成数据生成有效提升策略性能,在长时序任务中成功率从58%提升至72%。
原文摘要 · Abstract (English)
Action-conditioned world models have emerged as a promising paradigm for robot learning, offering a scalable alternative to costly real-world experimentation by generating action-consistent video rollouts. However, persistent world modeling remains challenging in manipulation: frequent end-effector occlusions and rapid wrist-camera motion make the current observation insufficient for predicting future views, causing models to forget or hallucinate scene details seen in earlier frames. Existing memory retrieval strategies often fail to identify informative history in dynamic manipulation scenarios. To address this limitation, we propose Mem-World, a memory-augmented multi-view action-conditioned world model. At its core, we present W-VMem, a 4D wrist-view-centered surfel-indexed memory that anchors historical observations to temporally evolving surface elements. By explicitly modeling when and where scene elements are observed, W-VMem enables geometry-aware retrieval of relevant history frames conditioned on future actions. During generation, relevant history frames are selected via surfel-based rendering and scoring, providing informative and non-redundant context for prediction. Extensive experiments show that Mem-World generates persistent rollouts in complex manipulation scenarios, enables more reliable policy evaluation than Ctrl-World, improving the Pearson correlation with real-world performance by 14.5\%, and supports effective policy improvement through synthetic data generation, increasing success rates from 58\% to 72\% on long-horizon tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。