arXiv:2505.16565cs.CV2025-05被引 10

用单目视频生成立体视频,一次完成修复与优化。

M2SVid: End-to-End Inpainting and Refinement for Monocular-to-Stereo Video Conversion

  • 基于深度重投影生成右视图,再用扩散模型统一修复和优化
  • 用户测试中胜率是第二名的2.6倍,速度提升6倍
  • 适合影视制作、VR内容生成等需要快速立体化场景

我们针对单目转立体视频问题,提出一种端到端的修复与优化架构。该方法以输入左视图视频、通过深度重投影得到的扭曲右视图以及遮挡掩码作为条件输入,扩展Stable Video Diffusion(SVD)模型,生成高质量右视图视频。为有效利用邻帧信息进行修复,我们修改SVD中的注意力层,对遮挡区域像素计算全注意力。模型通过最小化图像空间损失,在无需迭代扩散步骤的情况下端到端训练,实现高质量生成。实验表明,该方法在用户评测中胜出次数是第二名的2.6倍,且推理速度提升6倍。

原文摘要 · Abstract (English)

We tackle the problem of monocular-to-stereo video conversion and propose a novel architecture for inpainting and refinement of the warped right view obtained by depth-based reprojection of the input left view. We extend the Stable Video Diffusion (SVD) model to utilize the input left video, the warped right video, and the disocclusion masks as conditioning input to generate a high-quality right camera view. In order to effectively exploit information from neighboring frames for inpainting, we modify the attention layers in SVD to compute full attention for discoccluded pixels. Our model is trained to generate the right view video in an end-to-end manner without iterative diffusion steps by minimizing image space losses to ensure high-quality generation. Our approach outperforms previous state-of-the-art methods, being ranked best 2.6x more often than the second-place method in a user study, while being 6x faster.

立体视频视频生成图像修复

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。