arXiv:2410.00262cs.CV2024-10被引 8

将单视角视频转为立体视频,无需预估深度图。

ImmersePro: End-to-End Stereo Video Synthesis Via Implicit Disparity Learning

  • 用隐式视差引导生成立体对,避免显式深度图误差。
  • 在700万对数据上训练,性能超越现有方法11.76%以上。
  • 适合做视频3D化、虚拟现实内容生成的研究者。

我们提出 extit{ImmersePro},一种将单视角视频转换为立体视频的端到端框架。该框架采用双分支结构,包含视差分支和上下文分支,并结合时空注意力机制处理视频数据。通过隐式视差引导,直接从视频序列生成立体对,无需依赖显式视差图,从而降低视差估计模型带来的误差。此外,我们构建了 YouTube-SBS 数据集,涵盖423个来自YouTube的立体视频,共包含超过700万对立体图像,是当前规模最大的立体视频数据集,可用于训练与评估立体视频生成模型。实验表明, extit{ImmersePro} 在生成高质量立体视频方面表现优异,相较于最佳现有方法(stereo-from-mono),在L1、SSIM、PSNR指标上分别提升11.76%、6.39%、5.10%。

原文摘要 · Abstract (English)

We introduce \textit{ImmersePro}, an innovative framework specifically designed to transform single-view videos into stereo videos. This framework utilizes a novel dual-branch architecture comprising a disparity branch and a context branch on video data by leveraging spatial-temporal attention mechanisms. \textit{ImmersePro} employs implicit disparity guidance, enabling the generation of stereo pairs from video sequences without the need for explicit disparity maps, thus reducing potential errors associated with disparity estimation models. In addition to the technical advancements, we introduce the YouTube-SBS dataset, a comprehensive collection of 423 stereo videos sourced from YouTube. This dataset is unprecedented in its scale, featuring over 7 million stereo pairs, and is designed to facilitate training and benchmarking of stereo video generation models. Our experiments demonstrate the effectiveness of \textit{ImmersePro} in producing high-quality stereo videos, offering significant improvements over existing methods. Compared to the best competitor stereo-from-mono we quantitatively improve the results by 11.76\% (L1), 6.39\% (SSIM), and 5.10\% (PSNR).

立体视频隐式视差视频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。