让视频生成模型学会真实落物,靠小样本模拟数据微调+新奖励机制。
PISA Experiments: Exploring Physics Post-Training for Video Diffusion Models by Watching Stuff Drop

- 用少量模拟落物视频微调,提升模型物理准确性。
- 引入新奖励机制后,落物轨迹更符合真实物理规律。
- 适合关注视频模型物理可信度的开发者与研究者。
大规模预训练视频生成模型在内容创作上表现优异,但默认状态下无法可靠模拟真实物理世界。本文以物体自由下落这一基础物理任务为切入点,研究如何通过后训练提升模型的物理建模能力。实验表明,当前最先进的视频生成模型虽视觉效果逼真,却难以准确模拟落物行为。通过在少量模拟视频上微调,可有效诱导模型学习下落特征;进一步结合我们提出的新型奖励建模方法,性能显著提升。研究还揭示了后训练在泛化能力和分布建模上的关键局限。此外,本文发布了该任务的基准数据集,可用于追踪大规模视频生成模型的物理准确性进展。
原文摘要 · Abstract (English)
Large-scale pre-trained video generation models excel in content creation but are not reliable as physically accurate world simulators out of the box. This work studies the process of post-training these models for accurate world modeling through the lens of the simple, yet fundamental, physics task of modeling object freefall. We show state-of-the-art video generation models struggle with this basic task, despite their visually impressive outputs. To remedy this problem, we find that fine-tuning on a relatively small amount of simulated videos is effective in inducing the dropping behavior in the model, and we can further improve results through a novel reward modeling procedure we introduce. Our study also reveals key limitations of post-training in generalization and distribution modeling. Additionally, we release a benchmark for this task that may serve as a useful diagnostic tool for tracking physical accuracy in large-scale video generative model development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。