让视频生成模型能精准控制物理参数,提升真实感与可解释性。
Learning Explicit Physical Parameter Control and Benchmarking for Video Generation

- 基于物理注意力机制,显式绑定力、质量等参数到物体实例
- 在13万条物理模拟视频上训练,实现运动与物理定律高度一致
- 适合需要精确控制物理行为的场景生成与仿真研究
图像到视频生成的最新进展提升了视觉真实性,但其动态仍难以由明确的物理因素驱动,且多物体交互中实例约束易泄露或纠缠。我们指出两大缺失:大规模细粒度物理参数化数据,以及正确关联物理属性与实例的模型设计。为此,我们构建了包含13万条物理模拟视频的PhyParam-Dataset,覆盖五类刚体运动的力向量、材料属性与环境常数。在此基础上,提出PhyParam模型,通过轻量级物理注意力路由机制,显式引入物体受力、质量、摩擦、弹性和场景重力,并结合语义结构特征空间监督强化运动学习。同时建立PhyParam-Bench基准,从时序动态、空间稳定性与语义-物理对齐三个层面评估生成视频的物理一致性。实验表明,PhyParam在保持高视觉保真度的同时显著提升物理一致性,推动了图像到视频生成中的显式刚体物理参数控制。数据集、基准与代码将公开发布。
原文摘要 · Abstract (English)
Recent advances in image-to-video generation have improved visual realism, making physically grounded and controllable dynamics an important step toward future world simulation. Current models often generate plausible motion, but it is not reliably governed by explicit physical causes, and instance-level constraints can leak or become entangled in multi-object interactions. We attribute this gap to two missing pieces: large-scale, fine-grained physical parameterization, and model designs that correctly bind physical attributes to instances and emphasize dynamics over appearance. To bridge this gap, we introduce PhyParam-Dataset, an interaction-centric collection of 130K physically simulated videos with dense physical parameterization, including force vectors, object material properties, and environmental constants across five representative rigid-body motion types. Built on this data, we present PhyParam, a physics-guided image-to-video diffusion model that conditions on object-level forces, masses, friction, restitution, and scene-level gravity via a lightweight physical-attention routing mechanism, and further improves motion learning with semantic-structural feature-space supervision. We also establish PhyParam-Bench, a benchmark for physical-law consistency in image-to-video generation, with a multi-level protocol evaluating temporal dynamics, spatial stability, and semantic--physical alignment. Experiments show that PhyParam improves physical consistency while maintaining high visual fidelity, advancing explicit rigid-body physical-parameter control for image-to-video generation. We will publicly release the dataset, benchmark, and code to support future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。