无需训练即可生成符合物理规律的视频,通过分阶段推理实现精准控制。
PhyRPR: Training-Free Physics-Constrained Video Generation
- 分三阶段设计:先推理物理状态,再规划粗略运动,最后融合细化视觉。
- 在物理约束下生成的视频更合理,运动可控性显著提升。
- 适合需要精确物理模拟的场景,如科学可视化或虚拟实验。
基于扩散模型的视频生成方法虽能生成视觉逼真的视频,但常难以满足物理约束。主要原因是现有方法多为单阶段设计,将高层物理理解与低层视觉合成耦合,难以生成需显式物理推理的内容。为此,我们提出一种无需训练的三阶段流程——PhyRPR(PhyReason–PhyPlan–PhyRefine),将物理理解与视觉合成解耦。其中,PhyReason利用大模态模型进行物理状态推理,并用图像生成器合成关键帧;PhyPlan确定性地生成可控制的粗略运动骨架;PhyRefine通过潜在空间融合策略将该骨架注入扩散采样过程,以优化外观同时保持预定动态。这种分阶段设计实现了生成过程中的显式物理控制。大量物理约束下的实验表明,该方法在物理合理性与运动可控性方面均有显著提升。
原文摘要 · Abstract (English)
Recent diffusion-based video generation models can synthesize visually plausible videos, yet they often struggle to satisfy physical constraints. A key reason is that most existing approaches remain single-stage: they entangle high-level physical understanding with low-level visual synthesis, making it hard to generate content that require explicit physical reasoning. To address this limitation, we propose a training-free three-stage pipeline,\textit{PhyRPR}:\textit{Phy\uline{R}eason}--\textit{Phy\uline{P}lan}--\textit{Phy\uline{R}efine}, which decouples physical understanding from visual synthesis. Specifically, \textit{PhyReason} uses a large multimodal model for physical state reasoning and an image generator for keyframe synthesis; \textit{PhyPlan} deterministically synthesizes a controllable coarse motion scaffold; and \textit{PhyRefine} injects this scaffold into diffusion sampling via a latent fusion strategy to refine appearance while preserving the planned dynamics. This staged design enables explicit physical control during generation. Extensive experiments under physics constraints show that our method consistently improves physical plausibility and motion controllability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。