用视频扩散模型打通3D重建与生成的闭环,提升稀疏输入下的合成质量。
GenFusion: Closing the Loop between Reconstruction and Generation via Videos
- 基于重建驱动的视频扩散模型,用带伪影的RGB-D渲染图作为条件。
- 通过循环融合机制迭代加入生成修复帧,突破视角饱和限制。
- 在稀疏视图和缺失输入下仍能保持高质量视图合成,适合生成场景应用。
近期,3D重建与生成在新视角合成方面取得了显著进展,实现了高保真与高效率。然而,这两者之间存在明显的条件差距:可扩展的3D场景重建通常需要密集捕获的视图,而3D生成则多依赖单张或无输入视图,严重限制了实际应用。我们发现,这一现象的根源在于3D约束与生成先验之间的不匹配。为此,我们提出一种由重建驱动的视频扩散模型,学习以带有伪影的RGB-D渲染图为条件生成视频帧。此外,我们设计了一种循环融合流程,将生成模型输出的修复帧不断加入训练集,实现渐进式扩展,有效缓解了以往重建与生成流程中的视角饱和问题。我们的评估涵盖稀疏视图输入与遮挡输入下的视图合成任务,验证了方法的有效性。更多信息请见 https://genfusion.sibowu.com。
原文摘要 · Abstract (English)
Recently, 3D reconstruction and generation have demonstrated impressive novel view synthesis results, achieving high fidelity and efficiency. However, a notable conditioning gap can be observed between these two fields, e.g., scalable 3D scene reconstruction often requires densely captured views, whereas 3D generation typically relies on a single or no input view, which significantly limits their applications. We found that the source of this phenomenon lies in the misalignment between 3D constraints and generative priors. To address this problem, we propose a reconstruction-driven video diffusion model that learns to condition video frames on artifact-prone RGB-D renderings. Moreover, we propose a cyclical fusion pipeline that iteratively adds restoration frames from the generative model to the training set, enabling progressive expansion and addressing the viewpoint saturation limitations seen in previous reconstruction and generation pipelines. Our evaluation, including view synthesis from sparse view and masked input, validates the effectiveness of our approach. More details at https://genfusion.sibowu.com.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。