arXiv:2606.27374cs.ROcs.CV2026-06被引 1

用生成模型模拟旧任务,让机器人不存数据也能持续学习新技能。

World Action Models Enable Continual Imitation Learning with Recurrent Generative Replays

论文配图:World Action Models Enable Continual Imitation Learning with Recurrent Generative Replays
图 1 · 摘自论文原文
  • 用世界动作模型递归生成假回放轨迹,替代真实演示存储。
  • 实测可减少50%灾难性遗忘,接近需真实回放的最优方法。
  • 适合缺乏存储空间的机器人持续学习场景。

超越预测机器人动作,世界动作模型(WAMs)还能生成未来视觉观测。我们基于其生成能力提出循环生成回放(REGEN),一种无需存储原始人类示范即可持续模仿学习的框架。在持续适应过程中,REGEN仅依赖先前任务指令和当前任务观测,递归调用WAM生成伪回放轨迹供策略重练。仿真与真实机械臂实验表明,与顺序微调相比,REGEN将灾难性遗忘降低高达50%,性能接近需访问真实回放数据的特权方法。最后,我们分析生成回放的瓶颈,发现长时程视觉退化和动作-观测不一致是主要限制因素。结果证明WAMs是无需存储示范的持续机器人学习的有力基础。

原文摘要 · Abstract (English)

Going beyond predicting robot actions, World Action Models (WAMs) can also generate future visual observations. We build on this generative capability to propose Recurrent Generative Replay (REGEN), a continual imitation learning framework that synthesizes pseudo-replay trajectories, enabling a robot policy to rehearse previously learned tasks without storing their original human demonstrations. During continual adaptation, REGEN recursively queries the WAM to synthesize pseudo-replay trajectories conditioned only on prior task instructions and current-task observations. Experiments in both simulation and real-world manipulation settings show that REGEN reduces catastrophic forgetting by up to $50\%$ relative to sequential fine-tuning, while approaching the performance of privileged experience replay methods that require access to real replay data. Finally, we analyze the factors limiting generated replay, identifying long-horizon visual degradation and action-observation inconsistency as the primary bottlenecks. Our results establish WAMs as a promising foundation for continual robot learning without stored demonstrations.

持续学习生成回放机器人控制世界模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。