arXiv:2509.21309cs.CV2025-09中稿 · ICLR被引 39

用物理定律生成视频,让物体运动更真实可控。

NewtonGen: Physics-Consistent and Controllable Text-to-Video Generation via Neural Newtonian Dynamics

  • 引入可学习的牛顿动力学模型,约束视频生成过程
  • 支持不同初始条件下精确控制物体运动轨迹
  • 适合需要真实物理行为的视频生成场景

当前大规模文本到视频生成的主要瓶颈在于物理一致性和可控性。尽管已有进展,但最先进模型常生成不合理的运动,如物体向上掉落或速度方向突变。此外,这些模型缺乏对参数的精确控制,难以在不同初始条件下生成符合物理规律的动态。我们认为,这一根本限制源于现有模型仅从外观学习运动分布,而缺乏对底层动力学的理解。本文提出NewtonGen框架,融合数据驱动合成与可学习的物理原理。核心是可训练的神经牛顿动力学(NND),能建模和预测多种牛顿运动,从而将潜在的动力学约束注入视频生成过程。通过联合利用数据先验和动力学引导,NewtonGen实现物理一致的视频合成,并具备精确的参数控制能力。所有数据和代码已公开于https://github.com/pandayuanyu/NewtonGen。

原文摘要 · Abstract (English)

A primary bottleneck in large-scale text-to-video generation today is physical consistency and controllability. Despite recent advances, state-of-the-art models often produce unrealistic motions, such as objects falling upward, or abrupt changes in velocity and direction. Moreover, these models lack precise parameter control, struggling to generate physically consistent dynamics under different initial conditions. We argue that this fundamental limitation stems from current models learning motion distributions solely from appearance, while lacking an understanding of the underlying dynamics. In this work, we propose NewtonGen, a framework that integrates data-driven synthesis with learnable physical principles. At its core lies trainable Neural Newtonian Dynamics (NND), which can model and predict a variety of Newtonian motions, thereby injecting latent dynamical constraints into the video generation process. By jointly leveraging data priors and dynamical guidance, NewtonGen enables physically consistent video synthesis with precise parameter control. All data and code are available at https://github.com/pandayuanyu/NewtonGen

文本生成视频物理一致性可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。