分离物理推理与视觉生成,提升复杂场景视频生成的稳定性。
Motion Forcing: A Decoupled Framework for Robust Video Generation in Motion Dynamics
- 通过点-形状-外观分层框架解耦物理与视觉生成过程
- 在复杂交通场景中显著优于现有模型,保持视觉质量与物理一致性
- 适合需要高物理真实性的自动驾驶、机器人仿真等应用
视频生成的核心挑战在于同时实现高质量视觉效果、严格的物理一致性与精确可控性。现有模型在简单场景下表现良好,但在复杂场景(如碰撞或密集交通)中常失衡。为此,本文提出Motion Forcing框架,通过分层的‘点-形状-外观’范式显式解耦物理推理与视觉合成。该方法将生成过程分为三阶段:以稀疏几何锚点(点)建模动态,扩展为显式三维几何的动态深度图(形状),最后渲染高保真纹理(外观)。为增强物理理解,引入掩码点恢复策略:训练时随机遮蔽输入锚点,强制模型重建完整动态深度,从而学习惯性等潜在物理规律以推断缺失轨迹。在自动驾驶基准测试中,Motion Forcing显著超越现有先进模型,在复杂场景中维持三重平衡;物理与机器人任务评估进一步验证了其通用性。
原文摘要 · Abstract (English)
The ultimate goal of video generation is to satisfy a fundamental trilemma: achieving high visual quality, maintaining rigorous physical consistency, and enabling precise controllability. While recent models can maintain this balance in simple, isolated scenarios, we observe that this equilibrium is fragile and often breaks down as scene complexity increases (e.g., involving collisions or dense traffic). To address this, we introduce \textbf{Motion Forcing}, a framework designed to stabilize this trilemma even in complex generative tasks. Our key insight is to explicitly decouple physical reasoning from visual synthesis via a hierarchical \textbf{``Point-Shape-Appearance''} paradigm. This approach decomposes generation into verifiable stages: modeling complex dynamics as sparse geometric anchors (\textbf{Point}), expanding them into dynamic depth maps that explicitly resolve 3D geometry (\textbf{Shape}), and finally rendering high-fidelity textures (\textbf{Appearance}). Furthermore, to foster robust physical understanding, we employ a \textbf{Masked Point Recovery} strategy. By randomly masking input anchors during training and enforcing the reconstruction of complete dynamic depth, the model is compelled to move beyond passive pattern matching and learn latent physical laws (e.g., inertia) to infer missing trajectories. Extensive experiments on autonomous driving benchmarks show that Motion Forcing significantly outperforms state-of-the-art baselines, maintaining trilemma stability across complex scenes. Evaluations on physics and robotics further confirm our framework's generality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。