arXiv:2603.21017cs.RO2026-03

让机器人在意外干扰下仍能稳定执行任务,靠的是内建的想象预测能力。

Dreaming the Unseen: World Model-regularized Diffusion Policy for Out-of-Distribution Robustness

  • 用扩散世界模型与策略联合训练,共享视觉编码器提升状态预测能力
  • 在严重分布外扰动下成功率达73.8%(无想象时仅23.9%)
  • 可完全依赖内部想象运行,适合高风险或不可见环境下的机器人控制

扩散策略在视觉运动控制中表现优异,但在严重分布外(OOD)干扰下常失效,如物体意外移动或视觉退化。为此,我们提出梦想扩散策略(DDP),通过共享3D视觉编码器将扩散世界模型深度融入策略训练目标。这种联合优化使策略具备强鲁棒的状态预测能力。推理时遭遇突发OOD异常,DDP检测真实与想象之间的差异,主动舍弃受损视觉流,转而依赖自回归生成的隐空间动态进行内部‘想象’,生成安全轨迹后平稳回归现实。大量实验表明其卓越韧性:在MetaWorld上,分布外成功率高达73.8%(无预测想象时为23.9%);在严重空间位移下达83.3%(无想象时仅3.3%)。此外,在开环纯想象模式下仍保持76.7%的真实世界成功率,验证其强大适应性。

原文摘要 · Abstract (English)

Diffusion policies excel at visuomotor control but often fail catastrophically under severe out-of-distribution (OOD) disturbances, such as unexpected object displacements or visual corruptions. To address this vulnerability, we introduce the Dream Diffusion Policy (DDP), a framework that deeply integrates a diffusion world model into the policy's training objective via a shared 3D visual encoder. This co-optimization endows the policy with robust state-prediction capabilities. When encountering sudden OOD anomalies during inference, DDP detects the real-imagination discrepancy and actively abandons the corrupted visual stream. Instead, it relies on its internal "imagination" (autoregressively forecasted latent dynamics) to safely bypass the disruption, generating imagined trajectories before smoothly realigning with physical reality. Extensive evaluations demonstrate DDP's exceptional resilience. Notably, DDP achieves a 73.8% OOD success rate on MetaWorld (vs. 23.9% without predictive imagination) and an 83.3% success rate under severe real-world spatial shifts (vs. 3.3% without predictive imagination). Furthermore, as a stress test, DDP maintains a 76.7% real-world success rate even when relying entirely on open-loop imagination post-initialization.

扩散模型机器人控制泛化能力想象生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。