用深度扭曲与融合修复实现高质量立体视频转换
SpatialMe: Stereo Video Conversion Using Depth-Warping and Blend-Inpainting
- 通过多分支修复与掩码层级特征更新,提升立体图像生成质量
- 提出视差扩展策略,有效缓解前景溢出问题
- 构建1000段真实场景立体视频数据集,解决训练数据不足
立体视频转换旨在将单目视频转为沉浸式立体格式。尽管新视角合成技术取得进展,仍面临两大挑战:难以实现高保真且稳定的输出,以及高质量立体视频数据稀缺。本文提出SpatialMe框架,基于深度扭曲与融合修复技术。设计了基于掩码的层级特征更新(MHFU)精修模块,结合特征更新单元(FUU)与掩码机制,整合多分支修复模块输出。提出视差扩展策略,缓解前景溢出问题。此外,构建高质量真实世界立体视频数据集StereoV1K,包含1000段真实场景视频,分辨率为1180×1180,涵盖多种室内外场景。大量实验证明,该方法在生成立体视频方面优于现有最先进方法。
原文摘要 · Abstract (English)
Stereo video conversion aims to transform monocular videos into immersive stereo format. Despite the advancements in novel view synthesis, it still remains two major challenges: i) difficulty of achieving high-fidelity and stable results, and ii) insufficiency of high-quality stereo video data. In this paper, we introduce SpatialMe, a novel stereo video conversion framework based on depth-warping and blend-inpainting. Specifically, we propose a mask-based hierarchy feature update (MHFU) refiner, which integrate and refine the outputs from designed multi-branch inpainting module, using feature update unit (FUU) and mask mechanism. We also propose a disparity expansion strategy to address the problem of foreground bleeding. Furthermore, we conduct a high-quality real-world stereo video dataset -- StereoV1K, to alleviate the data shortage. It contains 1000 stereo videos captured in real-world at a resolution of 1180 x 1180, covering various indoor and outdoor scenes. Extensive experiments demonstrate the superiority of our approach in generating stereo videos over state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。