用扩散模型将单目视频转为高保真立体3D,适配VR设备。
StereoCrafter: Diffusion-based Generation of Long and High-fidelity Stereoscopic 3D from Monocular Videos
- 基于深度图和视频点云拼贴,生成立体视图。
- 在Stable Video Diffusion上微调,实现高质量立体补全。
- 支持长视频与多分辨率,适合苹果Vision Pro等设备。
本文提出一种新框架,将2D视频转换为沉浸式立体3D内容,以满足沉浸式体验中对3D内容日益增长的需求。利用基础模型作为先验知识,该方法克服了传统方法的局限性,显著提升性能,确保生成内容符合显示设备的高保真要求。系统包含两个核心步骤:基于深度图的视频点云拼贴(video splatting),用于图像扭曲与遮挡掩码提取;以及立体视频补全(stereo video inpainting)。采用预训练的Stable Video Diffusion作为主干网络,并引入针对立体视频补全任务的微调策略。为处理不同长度和分辨率的输入视频,探索了自回归策略与分块处理方法。最后,构建了一套复杂的数据处理流程,重建大规模、高质量数据集以支持训练。实验表明,该框架在2D到3D视频转换任务中表现优异,为苹果Vision Pro及3D显示器等设备提供了实用的沉浸式内容生成方案。
原文摘要 · Abstract (English)
This paper presents a novel framework for converting 2D videos to immersive stereoscopic 3D, addressing the growing demand for 3D content in immersive experience. Leveraging foundation models as priors, our approach overcomes the limitations of traditional methods and boosts the performance to ensure the high-fidelity generation required by the display devices. The proposed system consists of two main steps: depth-based video splatting for warping and extracting occlusion mask, and stereo video inpainting. We utilize pre-trained stable video diffusion as the backbone and introduce a fine-tuning protocol for the stereo video inpainting task. To handle input video with varying lengths and resolutions, we explore auto-regressive strategies and tiled processing. Finally, a sophisticated data processing pipeline has been developed to reconstruct a large-scale and high-quality dataset to support our training. Our framework demonstrates significant improvements in 2D-to-3D video conversion, offering a practical solution for creating immersive content for 3D devices like Apple Vision Pro and 3D displays. In summary, this work contributes to the field by presenting an effective method for generating high-quality stereoscopic videos from monocular input, potentially transforming how we experience digital media.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。