arXiv:2501.01409cs.CVcs.AI2025-01被引 7

首个实现3D一致性的视频生成模型,可同时生成真实视频与精准相机位姿。

JOG3R: Towards 3D-Consistent Video Generators

  • 联合训练视频生成与相机位姿估计任务,统一网络架构。
  • 在DUSt3R基准上达到竞争力的位姿估计性能,视频帧3D一致性显著提升。
  • 适合需要3D感知视频生成的研究者与应用开发者。

图像生成器的涌现能力催生了众多零样本或少样本应用。受此启发,我们探究视频生成器是否同样具备3D感知能力。以结构光恢复(structure-from-motion)作为3D感知任务,我们测试了视频生成器(以OpenSora为例)中间特征是否支持相机位姿估计。令人意外的是,初始仅发现两者间弱相关性。深入分析表明,尽管生成视频帧看似合理,但帧间缺乏真正的3D一致性。为此,我们提出联合训练策略,利用光照一致性生成与3D感知误差。研究发现,当前最优视频生成与相机位姿估计网络(如DUSt3R)具有共同结构,据此设计统一架构。所提模型名为\nameMethod,不仅能生成高质量视频,还能输出具备竞争力的相机位姿估计结果。综上,我们提出首个兼具3D一致性的视频生成模型,既生成真实视频,又可拓展至其他3D感知任务。

原文摘要 · Abstract (English)

Emergent capabilities of image generators have led to many impactful zero- or few-shot applications. Inspired by this success, we investigate whether video generators similarly exhibit 3D-awareness. Using structure-from-motion as a 3D-aware task, we test if intermediate features of a video generator - OpenSora in our case - can support camera pose estimation. Surprisingly, at first, we only find a weak correlation between the two tasks. Deeper investigation reveals that although the video generator produces plausible video frames, the frames themselves are not truly 3D-consistent. Instead, we propose to jointly train for the two tasks, using photometric generation and 3D aware errors. Specifically, we find that SoTA video generation and camera pose estimation (i.e.,DUSt3R [79]) networks share common structures, and propose an architecture that unifies the two. The proposed unified model, named \nameMethod, produces camera pose estimates with competitive quality while producing 3D-consistent videos. In summary, we propose the first unified video generator that is 3D-consistent, generates realistic video frames, and can potentially be repurposed for other 3D-aware tasks.

视频生成3D一致性位姿估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。