arXiv:2509.25191cs.CV2025-09被引 10

让3D基础模型直接生成密集新视角,突破传统依赖稀疏数据的瓶颈。

VGGT-X: When VGGT Meets Dense Novel View Synthesis

  • 用轻量VGGT处理上千张图像,降低显存压力
  • 自适应全局对齐提升输出质量,减少训练敏感性
  • 无需COLMAP,实现顶尖密集新视图合成效果

我们研究如何将3D基础模型(3DFMs)应用于密集新视图合成(NVS)。尽管基于NeRF和3DGS的新视图合成取得进展,现有方法仍依赖从结构光重建(SfM)获取的精确3D属性(如相机位姿和点云),而该过程在低纹理或低重叠场景中速度慢且易出错。近期3DFMs展现出比传统流程快数个数量级的潜力,并具备在线NVS能力。但多数验证仍局限于稀疏视角设置。我们的研究表明,将3DFMs直接扩展到密集视角面临两大根本障碍:显存消耗剧增,以及输出不完善导致初始化敏感的3D训练退化。为此,我们提出VGGT-X,包含内存高效的VGGT实现(支持1,000+图像)、自适应全局对齐以增强输出,并采用鲁棒的3DGS训练策略。大量实验表明,这些措施显著缩小了与COLMAP初始化方案的保真度差距,在无COLMAP的密集NVS和位姿估计上达到当前最优。此外,我们分析了与COLMAP初始化渲染之间仍存在的差距成因,为未来3D基础模型与密集NVS的发展提供洞见。

原文摘要 · Abstract (English)

We study the problem of applying 3D Foundation Models (3DFMs) to dense Novel View Synthesis (NVS). Despite significant progress in Novel View Synthesis powered by NeRF and 3DGS, current approaches remain reliant on accurate 3D attributes (e.g., camera poses and point clouds) acquired from Structure-from-Motion (SfM), which is often slow and fragile in low-texture or low-overlap captures. Recent 3DFMs showcase orders of magnitude speedup over the traditional pipeline and great potential for online NVS. But most of the validation and conclusions are confined to sparse-view settings. Our study reveals that naively scaling 3DFMs to dense views encounters two fundamental barriers: dramatically increasing VRAM burden and imperfect outputs that degrade initialization-sensitive 3D training. To address these barriers, we introduce VGGT-X, incorporating a memory-efficient VGGT implementation that scales to 1,000+ images, an adaptive global alignment for VGGT output enhancement, and robust 3DGS training practices. Extensive experiments show that these measures substantially close the fidelity gap with COLMAP-initialized pipelines, achieving state-of-the-art results in dense COLMAP-free NVS and pose estimation. Additionally, we analyze the causes of remaining gaps with COLMAP-initialized rendering, providing insights for the future development of 3D foundation models and dense NVS. Our project page is available at https://dekuliutesla.github.io/vggt-x.github.io/

3D生成新视图合成基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。