arXiv:2508.08048cs.CV2025-08TPAMI

用现有视频模型生成沉浸式3D视频,无需训练和姿态信息。

S^2VG: 3D Stereoscopic and Spatial Video Generation via Denoising Frame Matrix

  • 通过估计深度将单目视频映射到多视角,再用帧矩阵修复缺失内容。
  • 在Sora、Lumiere等模型上测试,生成的3D视频质量显著优于已有方法。
  • 适合想快速生成3D/空间视频的开发者和内容创作者使用。

尽管视频生成模型能高质量生成单目视频,但为沉浸式应用生成3D立体与空间视频仍是未充分探索的挑战。本文提出一种无需姿态、无需训练的方法,利用现成的单目视频生成模型生成沉浸式3D视频。该方法首先基于估计的深度信息,将生成的单目视频扭曲到预设相机视角;随后引入一种新颖的帧矩阵图像修复框架,利用原始视频生成模型合成不同视角与时间戳下的缺失内容,确保空间与时间一致性,且无需额外微调。此外,我们设计了双更新机制,缓解潜在空间中遮挡区域传播的负面效应,提升修复质量。最终生成的多视角视频可转化为立体对或优化为4D高斯用于空间视频合成。我们在Sora、Lumiere、WALT和Zeroscope等生成模型的视频上进行了实验,结果表明本方法显著优于先前方法。

原文摘要 · Abstract (English)

While video generation models excel at producing high-quality monocular videos, generating 3D stereoscopic and spatial videos for immersive applications remains an underexplored challenge. We present a pose-free and training-free method that leverages an off-the-shelf monocular video generation model to produce immersive 3D videos. Our approach first warps the generated monocular video into pre-defined camera viewpoints using estimated depth information, then applies a novel \textit{frame matrix} inpainting framework. This framework utilizes the original video generation model to synthesize missing content across different viewpoints and timestamps, ensuring spatial and temporal consistency without requiring additional model fine-tuning. Moreover, we develop a \dualupdate~scheme that further improves the quality of video inpainting by alleviating the negative effects propagated from disoccluded areas in the latent space. The resulting multi-view videos are then adapted into stereoscopic pairs or optimized into 4D Gaussians for spatial video synthesis. We validate the efficacy of our proposed method by conducting experiments on videos from various generative models, such as Sora, Lumiere, WALT, and Zeroscope. The experiments demonstrate that our method has a significant improvement over previous methods. Project page at: https://daipengwa.github.io/S-2VG_ProjectPage/

3D视频视频生成深度估计多视角

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。