无需深度估计,一键生成可调节立体感的高清3D视频
Elastic3D: Controllable Stereo Video Conversion with Guided Latent Decoding
- 基于条件扩散模型,用引导潜空间解码实现端到端转换
- 在三个真实数据集上超越现有方法,立体效果更清晰一致
- 用户可通过旋钮实时调节立体感强弱,操作直观
随着沉浸式3D内容需求增长,自动化单目视频转双目视频成为迫切需求。本文提出Elastic3D,一种可控、端到端的直接转换方法,将普通视频升级为双目视频。该方法基于(条件)潜空间扩散模型,避免了显式深度估计与图像扭曲带来的伪影。其高质量输出的关键在于一种新型引导式VAE解码器,确保生成结果既清晰又符合对极一致性。此外,用户可在推理时通过一个直观的标量旋钮,自由调节立体效果强度(即视差范围)。在三个真实世界双目视频数据集上的实验表明,该方法优于传统基于扭曲的方法和近期无扭曲基线,建立了可靠、可控双目视频转换的新标准。更多视频示例请见项目页:https://elastic3d.github.io。
原文摘要 · Abstract (English)
The growing demand for immersive 3D content calls for automated monocular-to-stereo video conversion. We present Elastic3D, a controllable, direct end-to-end method for upgrading a conventional video to a binocular one. Our approach, based on (conditional) latent diffusion, avoids artifacts due to explicit depth estimation and warping. The key to its high-quality stereo video output is a novel, guided VAE decoder that ensures sharp and epipolar-consistent stereo video output. Moreover, our method gives the user control over the strength of the stereo effect (more precisely, the disparity range) at inference time, via an intuitive, scalar tuning knob. Experiments on three different datasets of real-world stereo videos show that our method outperforms both traditional warping-based and recent warping-free baselines and sets a new standard for reliable, controllable stereo video conversion. Please check the project page for the video samples https://elastic3d.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。