arXiv:2509.20358cs.CV2025-09NeurIPS被引 45

让视频生成具备物理真实感和可控力场,实现逼真动态模拟

PhysCtrl: Generative Physics for Controllable and Physics-Grounded Video Generation

  • 用扩散模型学习四类材料的物理动态分布,支持力与参数控制
  • 在55万条仿真动画上训练,生成轨迹符合物理规律且可驱动视频生成
  • 适合需要真实物理行为的视频生成场景,如影视特效与虚拟仿真

现有视频生成模型虽能从文本或图像生成逼真视频,但常缺乏物理合理性与三维可控性。为克服此问题,我们提出PhysCtrl框架,实现基于物理参数与受力控制的图像到视频生成。核心是生成式物理网络,通过条件扩散模型学习弹性、沙粒、橡皮泥和刚性四类材料的物理动态分布。将物理动态表示为3D点轨迹,在由物理引擎生成的55万条动画构成的大规模合成数据集上进行训练。引入新型时空注意力模块以模拟粒子相互作用,并在训练中嵌入物理约束,确保生成结果的物理合理性。实验表明,PhysCtrl生成的运动轨迹具高度真实性,用于驱动图像到视频模型后,生成视频在视觉质量和物理可信度上均优于现有方法。

原文摘要 · Abstract (English)

Existing video generation models excel at producing photo-realistic videos from text or images, but often lack physical plausibility and 3D controllability. To overcome these limitations, we introduce PhysCtrl, a novel framework for physics-grounded image-to-video generation with physical parameters and force control. At its core is a generative physics network that learns the distribution of physical dynamics across four materials (elastic, sand, plasticine, and rigid) via a diffusion model conditioned on physics parameters and applied forces. We represent physical dynamics as 3D point trajectories and train on a large-scale synthetic dataset of 550K animations generated by physics simulators. We enhance the diffusion model with a novel spatiotemporal attention block that emulates particle interactions and incorporates physics-based constraints during training to enforce physical plausibility. Experiments show that PhysCtrl generates realistic, physics-grounded motion trajectories which, when used to drive image-to-video models, yield high-fidelity, controllable videos that outperform existing methods in both visual quality and physical plausibility. Project Page: https://cwchenwang.github.io/physctrl

视频生成物理模拟扩散模型可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。