arXiv:2512.00425cs.CV2025-12被引 19

让视频生成符合牛顿定律,用可验证奖励提升物理真实性

What about gravity in video generation? Post-Training Newton's Laws with Verifiable Rewards

  • 用光流和外观特征做物理代理,量化速度与质量
  • 引入运动约束与质量守恒双奖励,强制符合牛顿力学
  • 在6万样本基准上显著提升物理合理性与运动连贯性

近期视频扩散模型虽能生成视觉逼真的片段,但常违背基本物理规律——物体悬浮、加速度漂移、碰撞行为不一致,暴露出视觉真实与物理真实之间的差距。我们提出首个基于可验证奖励的物理引导后训练框架 $ exttt{NewtonRewards}$。不依赖人工或VLM反馈,该方法通过冻结的工具模型从生成视频中提取可测量代理:光流作为速度代理,高层外观特征作为质量代理。据此设计两种互补奖励:牛顿运动学约束(强制恒定加速度动态)与质量守恒奖励(防止退化解)。我们在新构建的大规模基准 $ exttt{NewtonBench-60K}$ 上评估了五类牛顿运动基元(自由落体、水平/抛物线投掷、滑下/上斜坡),结果表明,在所有基元的视觉与物理指标上,$ exttt{NewtonRewards}$ 均显著优于现有后训练方法,提升物理合理性、运动平滑性与时间连贯性,并在高度、速度、摩擦等分布外变化下保持稳健。结果证明,基于物理的可验证奖励为实现物理感知视频生成提供了可扩展路径。

原文摘要 · Abstract (English)

Recent video diffusion models can synthesize visually compelling clips, yet often violate basic physical laws-objects float, accelerations drift, and collisions behave inconsistently-revealing a persistent gap between visual realism and physical realism. We propose $\texttt{NewtonRewards}$, the first physics-grounded post-training framework for video generation based on $\textit{verifiable rewards}$. Instead of relying on human or VLM feedback, $\texttt{NewtonRewards}$ extracts $\textit{measurable proxies}$ from generated videos using frozen utility models: optical flow serves as a proxy for velocity, while high-level appearance features serve as a proxy for mass. These proxies enable explicit enforcement of Newtonian structure through two complementary rewards: a Newtonian kinematic constraint enforcing constant-acceleration dynamics, and a mass conservation reward preventing trivial, degenerate solutions. We evaluate $\texttt{NewtonRewards}$ on five Newtonian Motion Primitives (free fall, horizontal/parabolic throw, and ramp sliding down/up) using our newly constructed large-scale benchmark, $\texttt{NewtonBench-60K}$. Across all primitives in visual and physics metrics, $\texttt{NewtonRewards}$ consistently improves physical plausibility, motion smoothness, and temporal coherence over prior post-training methods. It further maintains strong performance under out-of-distribution shifts in height, speed, and friction. Our results show that physics-grounded verifiable rewards offer a scalable path toward physics-aware video generation.

视频生成物理建模扩散模型后训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。