arXiv:2505.18151cs.GRcs.AI2025-05ICCV被引 47

单图+动作指令生成动态3D场景,支持布料、液体等复杂物理效果。

WonderPlay: Dynamic 3D Scene Generation from a Single Image and Actions

  • 融合物理引擎与扩散模型,分步生成精细动态
  • 可生成布料、沙、雪、液体等10类以上动态效果
  • 用户通过动作指令即可控制复杂物理场景

WonderPlay 是一种新框架,将物理模拟与视频生成结合,从单张图像和动作指令生成动态3D场景。相比以往仅支持刚体或简单弹性运动的工作,WonderPlay 采用混合生成式模拟器,先用物理求解器生成粗粒度3D动态,再以该结果指导视频生成器输出更细腻逼真的运动序列。生成的视频反过来更新模拟场景,形成闭环。该方法融合了物理模拟的准确性与扩散模型的表达力,实现对布料、沙、雪、液体、烟雾、弹性体及刚体等多样内容的动态生成,且仅需单图输入。代码将公开。项目主页:https://kyleleey.github.io/WonderPlay/

原文摘要 · Abstract (English)

WonderPlay is a novel framework integrating physics simulation with video generation for generating action-conditioned dynamic 3D scenes from a single image. While prior works are restricted to rigid body or simple elastic dynamics, WonderPlay features a hybrid generative simulator to synthesize a wide range of 3D dynamics. The hybrid generative simulator first uses a physics solver to simulate coarse 3D dynamics, which subsequently conditions a video generator to produce a video with finer, more realistic motion. The generated video is then used to update the simulated dynamic 3D scene, closing the loop between the physics solver and the video generator. This approach enables intuitive user control to be combined with the accurate dynamics of physics-based simulators and the expressivity of diffusion-based video generators. Experimental results demonstrate that WonderPlay enables users to interact with various scenes of diverse content, including cloth, sand, snow, liquid, smoke, elastic, and rigid bodies -- all using a single image input. Code will be made public. Project website: https://kyleleey.github.io/WonderPlay/

3D生成物理模拟视频生成单图控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。