arXiv:2501.18804cs.CVcs.LG2025-01CVPR被引 12

用扩散模型直接生成新视角图像和深度图,无需中间3D表示

Zero-Shot Novel View and Depth Synthesis with Multi-View Geometric Diffusion

  • 基于射线图条件的扩散模型,融合多视角空间信息
  • 在6000万样本上训练,实现新视角合成与深度估计的领先性能
  • 支持零样本泛化,适合需要快速生成3D内容的场景

当前从稀疏带姿态图像进行3D场景重建的方法通常依赖于神经场、体素网格或3D高斯等中间3D表示,以实现多视角一致的外观与几何。本文提出MVGD,一种基于扩散的架构,可直接从任意数量输入视图生成像素级的新视角图像和深度图。该方法利用射线图条件,将不同视角的空间信息融入视觉特征,并引导新视角图像与深度图的生成。其核心在于通过可学习的任务嵌入,实现图像与深度图的多任务联合生成。模型在超过6000万个多视图样本(来自公开数据集)上进行训练,并提出高效一致的学习策略。此外,我们设计了一种新方法:通过逐步微调小模型来高效训练大模型,展现出良好的扩展性。大量实验表明,该方法在多个新视角合成基准、多视图立体匹配及视频深度估计任务中均达到当前最优性能。

原文摘要 · Abstract (English)

Current methods for 3D scene reconstruction from sparse posed images employ intermediate 3D representations such as neural fields, voxel grids, or 3D Gaussians, to achieve multi-view consistent scene appearance and geometry. In this paper we introduce MVGD, a diffusion-based architecture capable of direct pixel-level generation of images and depth maps from novel viewpoints, given an arbitrary number of input views. Our method uses raymap conditioning to both augment visual features with spatial information from different viewpoints, as well as to guide the generation of images and depth maps from novel views. A key aspect of our approach is the multi-task generation of images and depth maps, using learnable task embeddings to guide the diffusion process towards specific modalities. We train this model on a collection of more than 60 million multi-view samples from publicly available datasets, and propose techniques to enable efficient and consistent learning in such diverse conditions. We also propose a novel strategy that enables the efficient training of larger models by incrementally fine-tuning smaller ones, with promising scaling behavior. Through extensive experiments, we report state-of-the-art results in multiple novel view synthesis benchmarks, as well as multi-view stereo and video depth estimation.

扩散模型新视角合成深度估计3D生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。