arXiv:2606.27677cs.ROcs.CV2026-06被引 1

通过多尺度记忆增强,提升机器人长时任务的视觉与动作预测能力

DIM-WAM: World-Action Modeling with Diverse Historical Event Memory

论文配图:DIM-WAM: World-Action Modeling with Diverse Historical Event Memory
图 1 · 摘自论文原文
  • 引入多尺度记忆机制,融合近期、历史与全局任务状态信息
  • 在RMBench上将任务成功率从28.4%提升至69.8%,真实任务平均阶段成功率达91.5%
  • 适合长期依赖任务的机器人操控,尤其关注任务进度与历史事件建模

世界-动作模型在联合预测未来视觉状态与动作方面展现出良好的机器人操作性能。然而,现有方法主要依赖短期历史和短时预测,难以应对依赖早期观察与任务进展的长时任务。此类任务需要有效利用包括近期局部上下文、跨阶段历史事件、即时未来动态及全局任务进度在内的互补时间信息。为解决长期遗忘和全局任务状态感知不足问题,我们提出DiM-WAM,一种增强记忆的世界-动作模型,整合多尺度历史上下文、局部未来动态与全局任务进度。记忆模块从真实观测中提取紧凑视觉事件信息,通过独立的相似性合并更新多个记忆库,并读取带有身份与时间嵌入的长期上下文以指导视频与动作去噪。进度监督目标进一步促使记忆标记不仅编码已完成的历史事件,还包含当前任务阶段及其对剩余任务的影响。在RMBench上,DiM-WAM将平均成功率从LingBot-VA的28.4%提升至69.8%,超过显式记忆基线Mem-0(42.0%)。在四个真实世界Franka任务中,平均阶段成功率从70.7%提升至91.5%,完整任务成功率从52.5%提升至80.0%。

原文摘要 · Abstract (English)

World-action models have shown promising robot-manipulation performance by jointly predicting future visual states and actions. However, existing methods mainly rely on short-term history and short-horizon future prediction, which is insufficient for long-horizon tasks whose correct execution depends on earlier observations and task progress. Such temporally dependent tasks require effective use of complementary temporal information, including recent local context, cross-stage historical events, immediate future dynamics, and global task progress. To address long-term forgetting and poor awareness of the global task state, we introduce DiM-WAM, a memory-augmented world-action model that integrates multi-scale historical context, local future dynamics, and global task progress. The memory extracts compact visual event information from real observations, updates multiple memory banks through independent similarity-based merging, and then reads the bank-identity- and time-embedded long-term context to condition video and action denoising. A progress-supervision objective further encourages memory tokens to encode not only completed historical events but also the current task stage and its implications for the remaining task. On RMBench, DiM-WAM raises average success from 28.4% with LingBot-VA to 69.8%, exceeding the explicit-memory Mem-0 baseline at 42.0%. On four real-world Franka tasks, it improves average stage success from 70.7% to 91.5% and full-task success from 52.5% to 80.0%. Project page: https://wangkai-casia.github.io/dim-wam.

机器人操控长时记忆世界模型任务进度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。