arXiv:2606.14162cs.CV2026-06被引 1

通过隐式几何特征联合建模,解决视频生成中的结构漂移问题。

VideoWeave: Unlocking Geometric Consistency in Video Generation via Joint Geometry-Video Modeling

论文配图:VideoWeave: Unlocking Geometric Consistency in Video Generation via Joint Geometry-Video Modeling
图 1 · 摘自论文原文
  • 用隐式几何特征替代显式重建,降低上游误差影响。
  • 在共享去噪空间中联合建模几何与视频潜变量,提升一致性。
  • 适用于追求高质量且结构稳定的文本/图像转视频任务。

大规模视频扩散模型常因无法保持时序3D结构,导致视角变化时出现几何漂移和不自然运动。现有方法通常依赖深度图、点云或3D结构等显式几何重建作为条件、监督或奖励信号,使生成器对上游几何管道的误差敏感。本文提出VideoWeave,一种潜空间后训练框架,利用隐式几何模型特征约束生成分布,提供更灵活、非刚性的引导方式,减轻几何重建误差的影响。具体地,VideoWeave将这些特征转化为几何潜变量,并与视频潜变量在共享去噪空间中联合建模,使几何信息在训练过程中塑造生成分布。为支持该过程,我们构建了包含80,000段视频的GeoVid-80K数据集,配有配对的外观与几何表示。在文本到视频及图像到视频生成任务上的实验表明,VideoWeave在保持强视觉质量的同时显著提升了几何一致性。

原文摘要 · Abstract (English)

Large-scale video diffusion models often fail to preserve 3D structure over time, causing geometric drift and implausible motion under viewpoint changes. Existing methods usually enforce geometric consistency by using explicit geometry reconstructions, such as depth maps, point clouds, or reconstructed 3D structures, to define conditions, supervision, or reward signals, making the generator sensitive to errors from upstream geometry pipelines. We propose VideoWeave, a latent-space post-training framework that uses implicit geometry-model features to constrain the generative distribution, providing a more flexible and non-rigid form of guidance that mitigates the impact of reconstruction errors from geometry models. Specifically, VideoWeave adapts these features into geometry latents and jointly models them with video latents in a shared denoising space, allowing geometry to shape the generative distribution during training. To support this process, we build GeoVid-80K, an 80K-video dataset with paired appearance and geometry representations. Experiments on text-to-video and image-to-video generation show that VideoWeave improves geometric coherence while preserving strong visual quality. VideoWeave project page at https://videoweave.github.io/

视频生成几何一致性扩散模型潜空间建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。