让视频生成更懂物理,聚焦关键变形区域提升真实感
DeforM: Reasoning-Guided Physics-Aware Video Generation via Spatial-Temporal Masking

- 用视觉语言模型引导识别物理关键区域,生成时空掩码
- 在变形场景中显著提升视频真实度与物理一致性
- 适合需要高物理保真的视频生成研究者使用
视频生成模型虽已达到高视觉质量,但在生成符合物理规律的视频方面仍存挑战。与可由显式轨迹或公式描述的刚体运动不同,复杂形变动态难以合成。我们发现,缺乏对动态区域的物理推理导致无关区域分散模型注意力,引发生成失败。本文提出DeforM,一种基于物理推理引导的图像到视频生成框架,通过聚焦物理关键区域提升生成效果。为此,我们设计了VLM引导的物理推理模块DeforM-Reason,用于识别目标物体并生成时空掩码。同时提供两种物理引导策略:DeforM-Free用于无训练机制分析,DeforM-Injection作为强训练型生成器。实验表明,DeforM在视觉质量与物理一致性上均优于基线模型,显著提升形变场景生成的真实感。
原文摘要 · Abstract (English)
Video generation models achieve high visual quality but often struggle to generate physics-aware videos. Unlike rigid-body motion, which can be described by explicit trajectories or formulas, complex deformation dynamics remain challenging to synthesize. We observe that a lack of physical reasoning for localizing dynamic areas allows irrelevant regions to dilute the model's attention, leading to generation failure. In this paper, we propose DeforM, a reasoning-guided image-to-video generation framework that directs the model's focus toward physics-critical regions. To reason about and localize these critical regions, we introduce a VLM-guided physical reasoning module, DeforM-Reason, to identify target objects and generate spatial-temporal masks. For physical guidance, we develop two alternative strategies: DeforM-Free for training-free mechanism analysis and DeforM-Injection as a powerful training-based generator. Experimental results demonstrate that DeforM improves the realism of generated deformation scenarios, outperforming baseline models in both visual quality and physical consistency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。