用深度图约束视频生成,让动作更自然、结构更稳定
GeoVideo: Introducing Geometric Regularization into Video Generation Model
- 在扩散模型中加入每帧深度预测,引入几何正则化
- 多视角几何损失使不同帧的深度图在3D空间对齐
- 适合关注视频真实感与结构一致性的研究者
近期视频生成技术借助扩散变换器模型已能合成高质量、视觉逼真的视频片段。然而,多数方法仅在2D像素空间操作,缺乏对3D结构的显式建模,常导致时间上不一致的几何形态、不合理运动及结构伪影。本文通过在潜空间扩散模型中引入逐帧深度预测,将几何正则化损失融入视频生成过程。选用深度作为几何表示,得益于其预测精度的提升及与基于图像的潜编码器的良好兼容性。具体地,为实现时间上的结构一致性,提出一种跨帧多视角几何损失,将各帧预测的深度图在共享3D坐标系中对齐。该方法弥合了外观生成与3D结构建模之间的差距,显著提升了时空连贯性、形状一致性与物理合理性。在多个数据集上的实验表明,本方法生成结果比现有基线更加稳定且几何一致。
原文摘要 · Abstract (English)
Recent advances in video generation have enabled the synthesis of high-quality and visually realistic clips using diffusion transformer models. However, most existing approaches operate purely in the 2D pixel space and lack explicit mechanisms for modeling 3D structures, often resulting in temporally inconsistent geometries, implausible motions, and structural artifacts. In this work, we introduce geometric regularization losses into video generation by augmenting latent diffusion models with per-frame depth prediction. We adopted depth as the geometric representation because of the great progress in depth prediction and its compatibility with image-based latent encoders. Specifically, to enforce structural consistency over time, we propose a multi-view geometric loss that aligns the predicted depth maps across frames within a shared 3D coordinate system. Our method bridges the gap between appearance generation and 3D structure modeling, leading to improved spatio-temporal coherence, shape consistency, and physical plausibility. Experiments across multiple datasets show that our approach produces significantly more stable and geometrically consistent results than existing baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。