arXiv:2606.00267cs.CVcs.AI2026-06

让机器人模型主动想象危险场景,提前发现策略漏洞。

StressDream: Steering Video World Models for Robust Policy Evaluation and Improvement

论文配图:StressDream: Steering Video World Models for Robust Policy Evaluation and Improvement
图 1 · 摘自论文原文
  • 用文本指令引导扩散模型生成高影响但合理的未来画面。
  • 通过双目标优化,确保想象既符合语义又不偏离真实分布。
  • 适合做自动驾驶和机械臂控制的鲁棒性评估与改进。

视频世界模型(WMs)可通过条件于自车动作的未来观测想象,为策略评估与改进提供支持。然而,传统方法依赖默认想象,常忽略高影响结果,除非采样量极大。为此,本文提出StressDream,通过优化基于扩散模型的世界模型初始噪声,在推理时引导想象向指定的高影响但合理的结果偏移。针对高维噪声优化难题,设计两个互补目标:利用视觉-语言模型的语义目标提供有意义梯度,识别生成视频中的特定事件;以及可避免噪声漂移至分布外的合理性目标。在自动驾驶与机器人操作任务中,StressDream成功引导想象实现文本指定的高影响事件(如任务失败),从而揭示策略中潜在风险,实现更稳健的策略评估与改进。视频演示见 https://junwon.me/StressDream/。

原文摘要 · Abstract (English)

Video world models (WMs) have shown promise for policy evaluation and improvement by imagining realistic future observations conditioned on ego-robot actions. While WMs can model distributions over futures, policy evaluation and improvement typically rely on nominal imaginations, which can miss high-impact outcomes of robot actions unless prohibitively many samples are drawn. To enable robust policy evaluation and improvement over WM imaginations, we propose StressDream, which steers imaginations toward high-impact yet plausible outcomes specified at inference time by optimizing the initial noise of diffusion-based WMs. However, optimizing high-dimensional noise is challenging: the optimization must reason about nuanced, scene-dependent target events in generated videos while avoiding out-of-distribution (OOD) noise that yields implausible imaginations. We address this with two complementary objectives: a semantic objective with a Vision-Language Model that provides informative gradients by reasoning about the generated video, and a plausibility objective that prevents the optimized noise from drifting OOD. With state-of-the-art video world models for autonomous driving and robotic manipulation, we show that StressDream effectively steers imaginations toward high-impact yet plausible outcomes specified by text at inference time, such as task failures, enabling robust policy evaluation and improvement by identifying actions whose plausible futures include undesirable outcomes. Video results are available at https://junwon.me/StressDream/.

世界模型策略评估生成引导机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。