arXiv:2512.16915cs.CV2025-12被引 3

用生成先验统一转换单目视频为立体视频,效果更好且更快。

StereoPilot: Learning Unified and Efficient Stereo Conversion via Generative Priors

  • 直接生成目标视图,不依赖深度图和迭代采样。
  • 在双格式立体视频上实现更高画质与更高效计算。
  • 适合需要快速高质量3D内容生成的开发者与创作者。

立体显示设备(如VR头显、3D影院)的快速发展带来了对高质量立体视频内容的持续需求。然而,制作3D视频成本高且复杂,而自动单目转立体视频受限于多阶段‘深度-变形-修复’(DWI)流程,存在误差传播、深度歧义及平行与汇聚式立体配置格式不一致的问题。为此,我们提出首个大规模统一的立体视频转换数据集UniStereo,覆盖两种立体格式,支持公平评估与鲁棒模型训练。基于该数据集,我们构建了StereoPilot——一种高效的前馈模型,无需显式深度图或迭代扩散采样,直接合成目标视图。模型配备可学习域切换器与循环一致性损失,能无缝适配不同立体格式,显著提升一致性。大量实验表明,StereoPilot在视觉保真度与计算效率上均显著优于现有最优方法。

原文摘要 · Abstract (English)

The rapid growth of stereoscopic displays, including VR headsets and 3D cinemas, has led to increasing demand for high-quality stereo video content. However, producing 3D videos remains costly and complex, while automatic Monocular-to-Stereo conversion is hindered by the limitations of the multi-stage ``Depth-Warp-Inpaint'' (DWI) pipeline. This paradigm suffers from error propagation, depth ambiguity, and format inconsistency between parallel and converged stereo configurations. To address these challenges, we introduce UniStereo, the first large-scale unified dataset for stereo video conversion, covering both stereo formats to enable fair benchmarking and robust model training. Building upon this dataset, we propose StereoPilot, an efficient feed-forward model that directly synthesizes the target view without relying on explicit depth maps or iterative diffusion sampling. Equipped with a learnable domain switcher and a cycle consistency loss, StereoPilot adapts seamlessly to different stereo formats and achieves improved consistency. Extensive experiments demonstrate that StereoPilot significantly outperforms state-of-the-art methods in both visual fidelity and computational efficiency. Project page: https://hit-perfect.github.io/StereoPilot/.

立体视频图像生成深度学习视频转换

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。