用几何互反性实现单目视频转立体视频的自监督学习
Geometric Reciprocity: Unlocking Self-Supervision for Stereoscopic Video Generation

- 通过几何互反定理,从单目视频中解析出视差遮挡区域
- 无需标注数据,在真实视频上实现比现有方法更优的立体生成效果
- 适合做3D沉浸式内容生成的研究者和开发者
单目转立体视频可为沉浸式3D体验生成立体内容。在现代基于深度图的渲染(DIBR)方法中,视差遮挡区域的修补是主要瓶颈。基于训练的方法虽质量高,但依赖稀缺的立体图像对或存在域差距的合成数据。本文提出首个基于循环一致性的自监督框架,仅需单目视频即可训练。核心贡献是几何互反定理(GRT):在最近邻DIBR映射下,目标视角合成时的遮挡掩码等于从目标回传至源视角时丢失像素的掩码,可直接从单目图像解析计算测试阶段的遮挡掩码。该机制实现了训练与测试的一致性,支持从无限量单目视频中进行自监督学习,并显著优于无训练及监督型前沿方法。
原文摘要 · Abstract (English)
Monocular-to-stereo conversion synthesizes stereoscopic content from 2D videos for immersive 3D experiences. In modern Depth-Image-Based Rendering (DIBR) approaches, stereo inpainting of disocclusions is the critical bottleneck. Training-based methods achieve superior quality but rely on scarce stereo pairs or synthetic data with domain gaps. We address this through the first self-supervised framework learning from monocular videos via cycle consistency. Our key contribution is the Geometric Reciprocity Theorem (GRT): under the nearest-neighbor DIBR formulation, the disocclusion mask when synthesizing a target view equals the mask of pixels lost when warping back from target to source, enabling analytical computation of test-time disocclusion masks directly from monocular images. This yields train-test consistency for the stated warping formulation, supporting self-supervised learning from unlimited monocular videos and substantial improvements over training-free and supervised state-of-the-art methods. Project page: https://visual-ai.github.io/grt/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。