arXiv:2507.07982cs.CVcs.AI2025-07被引 80

让视频生成模型学会理解3D结构,提升画面一致性。

Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling

  • 用几何基础模型引导中间特征,强制学习3D结构。
  • 在视图和动作条件任务中,显著提升视觉质量和3D一致性。
  • 适合需要真实世界物理一致性的视频生成研究者。

视频本质上是动态3D世界的2D投影。然而,我们分析发现,仅基于原始视频训练的视频扩散模型往往无法捕捉其学习表征中的有意义几何结构。为弥合视频扩散模型与物理世界固有3D特性之间的差距,我们提出几何强迫(Geometry Forcing),一种简单而有效的方法,促使视频扩散模型内化3D表示。核心思想是通过将模型中间表示与几何基础模型的特征对齐,引导其趋向几何感知结构。为此,我们引入两种互补对齐目标:角度对齐(Angular Alignment),通过余弦相似度强制方向一致性;尺度对齐(Scale Alignment),通过从归一化扩散表示回归几何特征以保留尺度信息。我们在相机视角条件和动作条件视频生成任务上评估该方法。实验结果表明,相比基线方法,本方法在视觉质量与3D一致性方面均有显著提升。

原文摘要 · Abstract (English)

Videos inherently represent 2D projections of a dynamic 3D world. However, our analysis suggests that video diffusion models trained solely on raw video data often fail to capture meaningful geometric-aware structure in their learned representations. To bridge the gap between video diffusion models and the underlying 3D nature of the physical world, we propose Geometry Forcing, a simple yet effective method that encourages video diffusion models to internalize 3D representations. Our key insight is to guide the model's intermediate representations toward geometry-aware structure by aligning them with features from a geometric foundation model. To this end, we introduce two complementary alignment objectives: Angular Alignment, which enforces directional consistency via cosine similarity, and Scale Alignment, which preserves scale-related information by regressing geometric features from normalized diffusion representations. We evaluate Geometry Forcing on both camera-view conditioned and action-conditioned video generation tasks. Experimental results demonstrate that our method substantially improves visual quality and 3D consistency over the baseline methods. Project page: https://GeometryForcing.github.io.

视频生成3D建模扩散模型几何约束

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。