不训练模型,推理异常行为并实时修正生成视频的物理合理性。
Enhancing Physical Plausibility in Video Generation by Reasoning the Implausibility
- 用轻量级物理推理生成违背物理的行为提示
- 通过同步归一化与解耦去噪,即时抑制不合理运动
- 无需训练即可提升物理真实性,适合即插即用
扩散模型能生成逼真视频,但现有方法依赖大规模图文数据隐式学习物理规律,成本高且仍易生成违反基本物理定律的不合理运动。本文提出一种免训练框架,在推理阶段显式推理异常性并引导生成避开此类内容。具体而言,构建轻量级物理感知推理流水线,生成故意包含物理违规行为的反事实提示;提出新型同步解耦引导(SDG)策略,通过同步方向归一化缓解滞后抑制问题,利用轨迹解耦去噪减轻累积轨迹偏差,确保不合理内容在去噪全程中被及时、一致地抑制。跨多个物理领域的实验表明,该方法显著提升物理保真度,同时保持图像真实感,且无需额外训练。消融实验验证了物理推理模块与SDG的互补有效性,二者分别对抑制不合理内容和整体物理合理性提升至关重要。这建立了一种新的可即插即用的物理感知视频生成范式。
原文摘要 · Abstract (English)
Diffusion models can generate realistic videos, but existing methods rely on implicitly learning physical reasoning from large-scale text-video datasets, which is costly, difficult to scale, and still prone to producing implausible motions that violate fundamental physical laws. We introduce a training-free framework that improves physical plausibility at inference time by explicitly reasoning about implausibility and guiding the generation away from it. Specifically, we employ a lightweight physics-aware reasoning pipeline to construct counterfactual prompts that deliberately encode physics-violating behaviors. Then, we propose a novel Synchronized Decoupled Guidance (SDG) strategy, which leverages these prompts through synchronized directional normalization to counteract lagged suppression and trajectory-decoupled denoising to mitigate cumulative trajectory bias, ensuring that implausible content is suppressed immediately and consistently throughout denoising. Experiments across different physical domains show that our approach substantially enhances physical fidelity while maintaining photorealism, despite requiring no additional training. Ablation studies confirm the complementary effectiveness of both the physics-aware reasoning component and SDG. In particular, the aforementioned two designs of SDG are also individually validated to contribute critically to the suppression of implausible content and the overall gains in physical plausibility. This establishes a new and plug-and-play physics-aware paradigm for video generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。