arXiv:2604.28169cs.CVcs.AI2026-04被引 3

让视频生成模型学会控制物理属性,实现更真实的物体运动。

PhyCo: Learning Controllable Physical Priors for Generative Motion

论文配图:PhyCo: Learning Controllable Physical Priors for Generative Motion
图 1 · 摘自论文原文
  • 用10万+仿真视频训练,让模型学习摩擦、弹性和形变等物理属性。
  • 在Physics-IQ评测中显著提升物理真实性,优于现有方法。
  • 适合需要真实物理行为的视频生成场景,如动画和模拟。

现代视频扩散模型在外观生成上表现优异,但在物理一致性方面仍存在缺陷:物体漂移、碰撞缺乏真实反弹、材料响应与属性不符。我们提出PhyCo框架,将连续、可解释且基于物理的控制引入视频生成。该方法包含三部分:(i) 构建一个超过10万条的逼真仿真视频数据集,系统性地变化摩擦、弹性、形变和受力;(ii) 使用以像素对齐物理属性图为条件的ControlNet,对预训练扩散模型进行物理监督微调;(iii) 通过微调后的视觉-语言模型(VLM)引导奖励优化,以针对性物理问题评估生成视频并提供可微反馈。该组合使生成模型能在不依赖模拟器或几何重建的情况下,通过调整物理属性实现可控且物理一致的输出。在Physics-IQ基准测试中,PhyCo显著提升物理真实性,人眼评估也证实对物理属性的控制更清晰、更忠实。结果表明,这是迈向通用物理一致性生成视频模型的一条可扩展路径。

原文摘要 · Abstract (English)

Modern video diffusion models excel at appearance synthesis but still struggle with physical consistency: objects drift, collisions lack realistic rebound, and material responses seldom match their underlying properties. We present PhyCo, a framework that introduces continuous, interpretable, and physically grounded control into video generation. Our approach integrates three key components: (i) a large-scale dataset of over 100K photorealistic simulation videos where friction, restitution, deformation, and force are systematically varied across diverse scenarios; (ii) physics-supervised fine-tuning of a pretrained diffusion model using a ControlNet conditioned on pixel-aligned physical property maps; and (iii) VLM-guided reward optimization, where a fine-tuned vision-language model evaluates generated videos with targeted physics queries and provides differentiable feedback. This combination enables a generative model to produce physically consistent and controllable outputs through variations in physical attributes-without any simulator or geometry reconstruction at inference. On the Physics-IQ benchmark, PhyCo significantly improves physical realism over strong baselines, and human studies confirm clearer and more faithful control over physical attributes. Our results demonstrate a scalable path toward physically consistent, controllable generative video models that generalize beyond synthetic training environments.

视频生成物理模拟扩散模型可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。