arXiv:2512.10959cs.CV2025-12被引 2

无需深度图,用视角条件生成立体图像。

StereoSpace: Depth-Free Synthesis of Stereo Geometry via End-to-End Diffusion in a Canonical Space

  • 在标准空间中通过视角条件引导扩散模型生成立体对应关系。
  • 在分层和非朗伯场景中表现更优,实现清晰视差与强鲁棒性。
  • 适合需要无深度依赖立体生成的视觉应用开发者。

我们提出 StereoSpace,一种基于扩散模型的单目到立体图像合成框架,通过视角条件建模几何,无需显式深度或变形。在标准校正空间中,条件信息引导生成器端到端推断对应关系并填充遮挡。为确保公平且无泄露评估,引入端到端测试协议,测试时不使用任何真值或代理几何估计。评估指标聚焦下游相关性:iSQoE 体现感知舒适度,MEt3R 反映几何一致性。StereoSpace 在 warp & inpaint、latent-warping 与 warped-conditioning 三类方法中均表现领先,尤其在分层及非朗伯场景下展现出清晰视差与强鲁棒性,确立视角条件扩散模型为可扩展的无深度立体生成方案。

原文摘要 · Abstract (English)

We introduce StereoSpace, a diffusion-based framework for monocular-to-stereo synthesis that models geometry purely through viewpoint conditioning, without explicit depth or warping. A canonical rectified space and the conditioning guide the generator to infer correspondences and fill disocclusions end-to-end. To ensure fair and leakage-free evaluation, we introduce an end-to-end protocol that excludes any ground truth or proxy geometry estimates at test time. The protocol emphasizes metrics reflecting downstream relevance: iSQoE for perceptual comfort and MEt3R for geometric consistency. StereoSpace surpasses other methods from the warp & inpaint, latent-warping, and warped-conditioning categories, achieving sharp parallax and strong robustness on layered and non-Lambertian scenes. This establishes viewpoint-conditioned diffusion as a scalable, depth-free solution for stereo generation.

立体生成扩散模型无深度视角条件

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。