arXiv:2512.09363cs.CV2025-12被引 2

用单目视频生成高质量立体视频,保持三维结构准确

StereoWorld: Geometry-Aware Monocular-to-Stereo Video Generation

  • 联合输入单目视频并引入几何感知正则化,确保3D结构一致
  • 在1100万帧数据上训练,生成高分辨率立体视频无明显伪影
  • 适合XR内容创作、影视特效制作等需要立体视觉的场景

随着XR设备的普及,高质量立体视频需求激增,但其制作成本高且易出现伪影。为此,我们提出StereoWorld,一个端到端框架,复用预训练视频生成模型实现高保真单目到立体视频生成。该框架在输入单目视频的同时,显式引入几何感知正则化以保证三维结构准确性,并采用时空分块策略实现高效高分辨率合成。为支持大规模训练与评估,我们构建了一个包含超过1100万帧的高清立体视频数据集,其视距对齐自然人眼瞳距(IPD)。大量实验表明,StereoWorld显著优于现有方法,在视觉保真度和几何一致性方面均有明显提升。

原文摘要 · Abstract (English)

The growing adoption of XR devices has fueled strong demand for high-quality stereo video, yet its production remains costly and artifact-prone. To address this challenge, we present StereoWorld, an end-to-end framework that repurposes a pretrained video generator for high-fidelity monocular-to-stereo video generation. Our framework jointly conditions the model on the monocular video input while explicitly supervising the generation with a geometry-aware regularization to ensure 3D structural fidelity. A spatio-temporal tiling scheme is further integrated to enable efficient, high-resolution synthesis. To enable large-scale training and evaluation, we curate a high-definition stereo video dataset containing over 11M frames aligned to natural human interpupillary distance (IPD). Extensive experiments demonstrate that StereoWorld substantially outperforms prior methods, generating stereo videos with superior visual fidelity and geometric consistency. The project webpage is available at https://ke-xing.github.io/StereoWorld/.

立体视频视频生成几何感知XR内容

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。