arXiv:2603.22275cs.CV2026-03被引 7

用几何模型特征空间提升多视角图像生成的精度与速度。

Repurposing Geometric Foundation Models for Multi-view Diffusion

  • 以几何基础模型的特征空间作为扩散模型的潜在表示
  • 在2D质量与3D一致性上优于VAE和RAE,训练提速4.4倍以上
  • 无需文本到图像预训练,仍可媲美顶尖方法

尽管生成潜在空间的进展推动了单图生成的发展,但新颖视图合成(NVS)的最佳潜在空间仍不明确。NVS需要跨视角的几何一致性生成,而现有方法通常在视图无关的VAE潜在空间中运行。本文提出几何潜在扩散(GLD),将几何基础模型的几何一致特征空间重用于多视角扩散模型。我们证明这些特征不仅支持高保真RGB重建,还编码强跨视角几何对应关系,为NVS提供了理想潜在空间。实验表明,GLD在2D图像质量和3D一致性指标上均优于VAE和RAE,同时训练速度比VAE潜在空间快4.4倍以上。值得注意的是,即使从零开始训练扩散模型且未使用大规模文本到图像预训练,GLD仍保持与当前最先进方法相当的性能。

原文摘要 · Abstract (English)

While recent advances in generative latent spaces have driven substantial progress in single-image generation, the optimal latent space for novel view synthesis (NVS) remains largely unexplored. In particular, NVS requires geometrically consistent generation across viewpoints, but existing approaches typically operate in a view-independent VAE latent space. In this paper, we propose Geometric Latent Diffusion (GLD), a framework that repurposes the geometrically consistent feature space of geometric foundation models as the latent space for multi-view diffusion. We show that these features not only support high-fidelity RGB reconstruction but also encode strong cross-view geometric correspondences, providing a well-suited latent space for NVS. Our experiments demonstrate that GLD outperforms both VAE and RAE on 2D image quality and 3D consistency metrics, while accelerating training by more than 4.4x compared to the VAE latent space. Notably, GLD remains competitive with state-of-the-art methods that leverage large-scale text-to-image pretraining, despite training its diffusion model from scratch without such generative pretraining.

多视角生成扩散模型几何建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。