提出可迭代修正的扩散模型,让机器人序列预测能跨时域反复优化。
Diffusion ReRoll: Revisable Denoising for Robotic Sequential Prediction

- 分区域选择性重加噪,支持跨时域迭代修正。
- 在长序列规划中成功率提升超20%,动作预测提升56.5%。
- 适合需要高精度、可修正的机器人控制与视频-动作联合建模任务。
我们提出 Diffusion ReRoll,一种基于扩散模型的机器人序列预测框架,支持跨时域的可修订去噪。现有扩散模型通常仅执行单次单调去噪。Diffusion ReRoll 则选择性地对局部稳定区域重新加噪,其余区域继续去噪,使重加噪区域可借助全局上下文再次优化。这种结构化重加噪机制实现跨时域迭代修订,使早期与后期片段可相互修正,同时保持局部一致性。我们在长时序规划、策略学习和统一视频-动作建模任务中评估该方法,对比全序列扩散与基于 Diffusion Forcing 的因果去噪。在 OGBench PointMaze 与 AntMaze 上,相对于 Diffusion Forcing,Diffusion ReRoll 在匹配引导规划中平均成功率提升 21%;相对于 Diffuser,目标补全任务中提升 23%。在 LIBERO-10 多任务基准上,动作预测相对 Diffusion Policy 平均成功率提升 56.5%,且适用于不同预测时长与历史长度。在统一视频-动作预测中,其策略与逆动力学性能更优,尤其在分布外评估下表现突出,且实现最佳动作-视频一致性。结果表明,结构化重加噪是可修订机器人序列生成的有效机制。
原文摘要 · Abstract (English)
We propose Diffusion ReRoll, a diffusion-based framework for robotic sequential prediction that enables revisable denoising over horizons. Existing diffusion-based sequence predictors typically perform a single monotonic denoising process. In contrast, Diffusion ReRoll selectively re-noises regions that have become locally stable while the remaining regions continue denoising, so the re-noised regions can be refined again using context from the rest of the horizon. This structured re-noising enables iterative cross-horizon revision, allowing earlier and later segments to revise one another, while maintaining local consistency. We evaluate Diffusion ReRoll against full-sequence diffusion and causal denoising based on Diffusion Forcing across long-horizon planning, policy learning, and unified video-action modeling. On OGBench PointMaze and AntMaze, Diffusion ReRoll achieves relative gains in average success rate of 21% over Diffusion Forcing in matched guidance-based planning and 23% over Diffuser in matched goal-inpainting. In diffusion-policy-style action prediction, Diffusion ReRoll improves average success by 56.5% relative to Diffusion Policy across different prediction horizons and history lengths on the LIBERO-10 multi-task benchmark. In unified video-action prediction, Diffusion ReRoll improves policy and inverse dynamics performance, especially under out-of-distribution evaluation, and achieves the best action-video consistency. These results support structured re-noising as an effective mechanism for revisable robotic sequence generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。