arXiv:2606.04811cs.CV2026-06

用机器人执行视频动作,测试生成模型是否懂物理。

Dream.exe: Can Video Generation Models Dream Executable Robot Manipulation?

论文配图:Dream.exe: Can Video Generation Models Dream Executable Robot Manipulation?
图 1 · 摘自论文原文
  • 让模型生成操作视频,转成机器人轨迹执行
  • 8个模型在101项任务中达成可执行成功率
  • 视觉质量好不等于能落地,适合评估模型真实物理理解

视频生成模型虽在视觉表现上进步显著,但其输出仍局限于虚拟空间。我们提出以机器人操作为检验标准:若模型真正理解物理规律,其生成的动作应能在现实世界中执行。为此,我们构建Dream.exe评估框架,通过视频到执行的流水线实现这一目标——给定场景图像和任务描述,模型生成操作视频,将其运动转化为机器人轨迹,并在物理模拟器中执行,从而获得纯视觉指标无法提供的真实反馈信号。我们在涵盖101项人工标注操作任务、分三个物理复杂度层级的基准上评估了8个模型(包括前沿闭源、开源生成模型及专用机器人模型)。结果表明,多个模型展现出可观测的执行成功率,说明互联网规模数据训练出的生成先验已蕴含有意义的物理知识。然而,视觉质量与可执行性无强关联,揭示了现有评估体系未覆盖的能力维度。Dream.exe将开源于https://github.com/showlab/Dream.exe。

原文摘要 · Abstract (English)

Video generation models have made impressive strides in synthesizing visually compelling content, yet their outputs remain confined to the virtual domain. A natural question follows: how well do these models reflect the physical world when their generated videos leave the screen and enter reality? We propose robotic manipulation as a concrete, measurable window onto this question: if a model has truly internalized physical laws, the motion it depicts should translate into executable robot behavior. We introduce Dream$.$exe, an evaluation framework that operationalizes this criterion through a video-to-execution pipeline. Given a scene image and a task description, Dream$.$exe synthesizes a manipulation video, converts the generated motion into robot trajectories, and executes them in a physics simulator, yielding a grounding signal that purely visual metrics cannot offer. Using this pipeline, we evaluate 8 models spanning frontier closed-source generators, open-source generators, and robot-specific models. Our benchmark covers 101 manually curated manipulation tasks at three levels of physical complexity, measured across visual quality, trajectory fidelity, and execution success. Encouragingly, several models achieve measurable execution success, suggesting that generative priors learned from internet-scale data already encode meaningful physical knowledge. Yet visual quality proves a poor predictor of executability, exposing a dimension of model capability that standard visual evaluations do not capture. Dream$.$exe will be open-sourced at https://github.com/showlab/Dream.exe.

视频生成机器人物理理解评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。