arXiv:2411.11934cs.CVcs.AI2024-11CVPR被引 11

用单目视频自监督生成立体视频,解决数据少与时序不一致难题。

SpatialDreamer: Self-supervised Stereo Video Synthesis from Monocular Input

  • 通过深度引导的视频生成模块,自动生成带几何和时间先验的成对视频。
  • 在多个基准上超越现有方法,立体偏差控制更优,时序一致性更强。
  • 适合做虚拟现实、空间计算中立体视频生成的研究者与开发者。

从单目输入合成立体视频是空间计算与虚拟现实中的高难度任务,主要挑战在于高质量成对立体视频数据不足,以及帧间时空一致性难以维持。现有方法多直接将新视角合成技术应用于视频,但存在难以有效表征动态场景、需大量训练数据等局限。本文提出一种基于视频扩散模型的自监督立体视频合成框架 SpatialDreamer。首先,为缓解数据不足问题,设计深度引导视频生成模块 DVG,采用前向-后向渲染机制生成具有几何与时间先验的成对视频。利用 DVG 生成的数据,提出 RefinerNet 与自监督训练框架,实现高效专一训练。更重要的是,设计一致性控制模块,包含立体偏差强度度量与时间交互学习模块(TIL),分别保障几何与时序一致性。在多个基准上的评估结果表明,该方法性能显著优于现有方法。

原文摘要 · Abstract (English)

Stereo video synthesis from a monocular input is a demanding task in the fields of spatial computing and virtual reality. The main challenges of this task lie on the insufficiency of high-quality paired stereo videos for training and the difficulty of maintaining the spatio-temporal consistency between frames. Existing methods primarily address these issues by directly applying novel view synthesis (NVS) techniques to video, while facing limitations such as the inability to effectively represent dynamic scenes and the requirement for large amounts of training data. In this paper, we introduce a novel self-supervised stereo video synthesis paradigm via a video diffusion model, termed SpatialDreamer, which meets the challenges head-on. Firstly, to address the stereo video data insufficiency, we propose a Depth based Video Generation module DVG, which employs a forward-backward rendering mechanism to generate paired videos with geometric and temporal priors. Leveraging data generated by DVG, we propose RefinerNet along with a self-supervised synthetic framework designed to facilitate efficient and dedicated training. More importantly, we devise a consistency control module, which consists of a metric of stereo deviation strength and a Temporal Interaction Learning module TIL for geometric and temporal consistency ensurance respectively. We evaluated the proposed method against various benchmark methods, with the results showcasing its superior performance.

立体视频自监督视频生成扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。