arXiv:2602.21929cs.CV2026-02中稿 · CVPR被引 6

用几何信息当上下文,让视频生成更稳更准。

Geometry-as-context: Modulating Explicit 3D in Scene-consistent Video Generation to Geometry Context

  • 以相机位姿为引导,用几何上下文迭代生成3D一致视频。
  • 在单向与往返轨迹上,场景一致性优于已有方法。
  • 适合需要精准相机控制的3D视频生成任务。

场景一致视频生成旨在根据相机轨迹生成探索3D场景的视频。以往方法依赖带外部记忆的生成模型,或通过迭代3D重建与修补,因中间输出错误、不可微过程及分离模型导致误差累积。为此,我们提出「几何即上下文」机制:基于自回归相机控制视频生成模型,迭代执行两步:(1) 估计当前视角所需几何信息以支持3D重建;(2) 模拟并恢复由3D场景渲染出的新视角图像。在此多任务框架下,我们设计了相机门控注意力模块,增强模型对相机姿态的利用能力。训练阶段,文本上下文决定生成几何或RGB图像;推理时,随机丢弃几何上下文,确保仅输出RGB图像。该方法在单向与往返轨迹的场景视频生成任务中表现优异,显著提升场景一致性和相机控制能力。

原文摘要 · Abstract (English)

Scene-consistent video generation aims to create videos that explore 3D scenes based on a camera trajectory. Previous methods rely on video generation models with external memory for consistency, or iterative 3D reconstruction and inpainting, which accumulate errors during inference due to incorrect intermediary outputs, non-differentiable processes, and separate models. To overcome these limitations, we introduce ``geometry-as-context". It iteratively completes the following steps using an autoregressive camera-controlled video generation model: (1) estimates the geometry of the current view necessary for 3D reconstruction, and (2) simulates and restores novel view images rendered by the 3D scene. Under this multi-task framework, we develop the camera gated attention module to enhance the model's capability to effectively leverage camera poses. During the training phase, text contexts are utilized to ascertain whether geometric or RGB images should be generated. To ensure that the model can generate RGB-only outputs during inference, the geometry context is randomly dropped from the interleaved text-image-geometry training sequence. The method has been tested on scene video generation with one-direction and forth-and-back trajectories. The results show its superiority over previous approaches in maintaining scene consistency and camera control.

3D生成视频生成相机控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。