通过四视角几何引导,生成更符合物理规律的视频。
OrthoPhys: Physically Plausible Video Generation with Orthogonal-View Geometry Guidance
- 先生成四个正交视角的同步视频,强制三维空间一致性。
- 在160K视频数据集上训练,显著提升运动真实感。
- 适合需要高物理真实性的视频生成任务。
视频生成领域虽在视觉质量上取得进展,但保持物理一致性仍是根本挑战。现实物体运动发生在三维空间,而视频仅提供视点相关的投影。为此,我们提出OrthoPhys,一种两阶段框架,利用四视角正交几何引导来确保物理合理性。第一阶段生成同步的四个正交视角前景动态视频,通过跨视角的几何增强注意力机制,有效约束三维空间一致性并隐式锚定运动属性。第二阶段以这些物理一致的正交前景为刚性引导,合成完整视频,学习前景与背景的交互关系。为支持该训练范式,我们构建了PhysMV数据集,包含40,000个场景,每个场景有四个正交视点,共160,000条视频序列。大量实验表明,OrthoPhys在物理真实性与时空连贯性上显著优于现有方法。
原文摘要 · Abstract (English)
Recent progress in video generation has led to substantial improvements in visual fidelity, yet ensuring physically consistent motion remains a fundamental challenge. Intuitively, this limitation can be attributed to the fact that real-world object motion unfolds in three-dimensional space, while video observations provide only partial, view-dependent projections of such dynamics. To address these issues, we propose OrthoPhys, a two-stage framework that leverages orthogonal-view geometry guidance to enforce physical plausibility. Instead of directly generating unstructured 2D videos, our first stage generates synchronized, four-view orthogonal videos of the foreground dynamics. By incorporating a geometry-enhanced attention mechanism across these orthogonal views, this stage effectively enforces 3D spatial coherence and implicitly grounds the motion in physical attributes. In the second stage, these physically consistent orthogonal foregrounds serve as rigid guidance to synthesize the final complete video, seamlessly learning the interaction between foreground dynamics and the background context. To support this orthogonal-view training paradigm, we construct PhysMV, a dataset containing 40K scenes, each consisting of four orthogonal viewpoints, resulting in a total of 160K video sequences. Extensive experiments demonstrate that OrthoPhys significantly improves physical realism and spatial-temporal coherence over existing video generation methods. Project page: https://anonymous.4open.science/w/Phys4D/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。