无需配对数据,用扩散模型生成高质量立体视频。
DissolveStereo: Coarse Depth Injection for Zero-Shot Stereo Video Generation
- 通过噪声重启初始化立体感知潜在表示,提升初始一致性。
- 使用溶解深度图降低高频信息,使时序更平滑,帧间不闪烁。
- 零样本训练,适合快速生成立体视频,提升视觉真实感。
生成高质量立体视频需要保持左右视图间的一致性与时间连贯性。尽管扩散模型在图像和视频生成方面取得进展,但维持左右视角间时空一致性的挑战依然存在。本文提出DissolveStereo,一种无需配对训练数据的零样本立体视频生成框架,利用视频扩散先验。关键创新包括:采用噪声重启策略初始化立体感知潜在表示,以及迭代精炼过程逐步调和潜在空间,缓解时间闪烁和视图不一致问题。我们提出使用溶解深度图,以减少高频深度信息,简化潜在空间操作。全面评估显示,该方法显著提升深度一致性与时序平滑性。在基线方法上,埃皮波尔一致性指标MEt3R提升11.7%。用户研究进一步表明,帧质量感知提高8.0%,时序连贯性感知提升10.9%。代码已开源:https://github.com/shijianjian/DissolveStereo。
原文摘要 · Abstract (English)
Generating high-quality stereo videos requires consistent depth perception and temporal coherence across frames. Despite advances in image and video synthesis using diffusion models, producing high-quality stereo videos remains a challenging task due to the difficulty of maintaining consistent temporal and spatial coherence between left and right views. We introduce DissolveStereo, a novel framework for zero-shot stereo video generation that leverages video diffusion priors without requiring paired training data. Our key innovations include a noisy restart strategy to initialize stereo-aware latent representations and an iterative refinement process that progressively harmonizes the latent space, addressing issues like temporal flickering and view inconsistencies. Importantly, we propose the use of dissolved depth maps to streamline latent space operations by reducing high-frequency depth information. Our comprehensive evaluations, including quantitative metrics and user studies, demonstrate that DissolveStereo produces high-quality stereo videos with enhanced depth consistency and temporal smoothness. In terms of epipolar consistency, our method achieves an 11.7% improvement in MEt3R score over the current state-of-the-art. Furthermore, user studies indicate strong perceptual gains over the previous arts, with an 8.0% higher perceived frame quality and 10.9% higher perceived temporal coherence. Our code is in https://github.com/shijianjian/DissolveStereo.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。