arXiv:2411.16157cs.CV2024-11CVPR被引 30

用3D先验让单图生成百张新视角,保持三维一致性。

MVGenMaster: Scaling Multi-View Generation from Any Image via 3D Priors Enhanced Diffusion Model

  • 基于3D先验与真实深度、相机位姿对齐,提升多视角生成效果。
  • 单次前向传播可生成最多100个新视角,支持任意参考视图。
  • 构建了含160万场景的大型数据集MvD-1M,推动方法规模化。

我们提出MVGenMaster,一种融合3D先验的多视角扩散模型,用于解决多样化的新型视图合成(NVS)任务。该模型利用通过度量深度和相机位姿映射的3D先验,显著提升了NVS中的泛化能力与三维一致性。其采用简单高效的流程,仅需一次前向传播即可根据可变参考视图和相机位姿生成最多100个新视角。此外,我们构建了一个大规模多视角图像数据集MvD-1M,包含高达160万场景,并配有精确对齐的度量深度,用于训练MVGenMaster。同时,我们提出了多种训练与模型改进策略,以适应大规模数据集。在跨域与非域基准上的广泛评估表明,所提方法与数据设计均具有效性。模型与代码将发布于https://github.com/ewrfcas/MVGenMaster/。

原文摘要 · Abstract (English)

We introduce MVGenMaster, a multi-view diffusion model enhanced with 3D priors to address versatile Novel View Synthesis (NVS) tasks. MVGenMaster leverages 3D priors that are warped using metric depth and camera poses, significantly enhancing both generalization and 3D consistency in NVS. Our model features a simple yet effective pipeline that can generate up to 100 novel views conditioned on variable reference views and camera poses with a single forward process. Additionally, we have developed a comprehensive large-scale multi-view image dataset called MvD-1M, comprising up to 1.6 million scenes, equipped with well-aligned metric depth to train MVGenMaster. Moreover, we present several training and model modifications to strengthen the model with scaled-up datasets. Extensive evaluations across in- and out-of-domain benchmarks demonstrate the effectiveness of our proposed method and data formulation. Models and codes will be released at https://github.com/ewrfcas/MVGenMaster/.

多视角生成扩散模型3D先验视图合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。