arXiv:2412.03517cs.CV2024-12CVPR被引 15

无需对齐就能用多张无姿态图生成高质量新视角。

NVComposer: Boosting Generative Novel View Synthesis with Multiple Sparse and Unposed Images

  • 用双流扩散模型同时生成新视角和相机位姿。
  • 在多个输入视图下合成质量显著提升,最高达4.2%的PSNR增益。
  • 适合无预对齐数据的通用图像生成场景,尤其适用于低重叠率场景。

近期生成模型在多视角数据的新生视角合成(NVS)上取得显著进展。然而,现有方法依赖外部多视角对齐流程,如显式位姿估计或预先重建,这限制了其灵活性与可及性,尤其当视角间重叠不足或存在遮挡时对齐不稳定。本文提出NVComposer,一种无需显式外部对齐的新方法。NVComposer通过引入两个关键组件,使生成模型能隐式推断多条件视角间的空间与几何关系:1)图像-位姿双流扩散模型,可同步生成目标新视角和条件相机位姿;2)几何感知特征对齐模块,在训练中从密集立体模型中提取几何先验。大量实验表明,NVComposer在生成式多视角NVS任务中达到当前最优性能,摆脱对外部对齐的依赖,提升模型可及性。随着输入无姿态视图数量增加,合成质量显著提升,凸显其在更灵活、易用的生成式NVS系统中的潜力。

原文摘要 · Abstract (English)

Recent advancements in generative models have significantly improved novel view synthesis (NVS) from multi-view data. However, existing methods depend on external multi-view alignment processes, such as explicit pose estimation or pre-reconstruction, which limits their flexibility and accessibility, especially when alignment is unstable due to insufficient overlap or occlusions between views. In this paper, we propose NVComposer, a novel approach that eliminates the need for explicit external alignment. NVComposer enables the generative model to implicitly infer spatial and geometric relationships between multiple conditional views by introducing two key components: 1) an image-pose dual-stream diffusion model that simultaneously generates target novel views and condition camera poses, and 2) a geometry-aware feature alignment module that distills geometric priors from dense stereo models during training. Extensive experiments demonstrate that NVComposer achieves state-of-the-art performance in generative multi-view NVS tasks, removing the reliance on external alignment and thus improving model accessibility. Our approach shows substantial improvements in synthesis quality as the number of unposed input views increases, highlighting its potential for more flexible and accessible generative NVS systems. Our project page is available at https://lg-li.github.io/project/nvcomposer

新视角合成扩散模型无对齐生成多视图重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。