arXiv:2506.09885cs.CV2025-06

数据越多越强:无3D结构和相机位姿也能实现顶尖新视角合成。

The Less You Depend, The More You Learn: Synthesizing Novel Views from Sparse, Unposed Images with Minimal 3D Knowledge

  • 摒弃3D结构与相机位姿,仅靠海量2D图像学习隐式三维感知。
  • 在大规模数据下性能超越依赖3D先验的方法,且随数据增长持续提升。
  • 适合追求高可扩展性、无标注数据场景的视觉生成研究者。

近年来,前馈式新视角合成(NVS)发展出两种设计范式:依赖显式3D知识的方法(如NeRF、3DGS)和基于数据的方法(从大规模图像中隐式学习3D结构)。本文通过全面分析发现,依赖更少3D知识的方法在训练数据增加时性能提升更快,最终超越依赖3D先验的方法,提出“越少依赖,学得越多”的规律。基于此,我们设计了一种无需显式场景结构或位姿标注的前馈NVS框架。该方法完全依赖大量2D图像学习隐式3D感知,训练与推理均无需任何位姿信息。大量实验表明,模型在无位姿条件下达到当前最优性能,甚至超越依赖已知位姿的方法,验证了数据驱动范式的有效性及可扩展性原则的指导价值。

原文摘要 · Abstract (English)

Recent advances in feed-forward Novel View Synthesis (NVS) have led to a divergence between two design philosophies: bias-driven methods, which rely on explicit 3D knowledge, such as handcrafted 3D representations (e.g., NeRF and 3DGS) and camera poses annotated by Structure-from-Motion algorithms, and data-centric methods, which learn to understand 3D structure implicitly from large-scale imagery data. This raises a fundamental question: which paradigm is more scalable in an era of ever-increasing data availability? In this work, we conduct a comprehensive analysis of existing methods and uncover a critical trend that the performance of methods requiring less 3D knowledge accelerates more as training data increases, eventually outperforming their 3D knowledge-driven counterparts, which we term "the less you depend, the more you learn." Guided by this finding, we design a feed-forward NVS framework that removes both explicit scene structure and pose annotation reliance. By eliminating these dependencies, our method leverages great scalability, learning implicit 3D awareness directly from vast quantities of 2D images, without any pose information for training or inference. Extensive experiments demonstrate that our model achieves state-of-the-art NVS performance, even outperforming methods relying on posed training data. The results validate not only the effectiveness of our data-centric paradigm but also the power of our scalability finding as a guiding principle.

新视角合成数据驱动无位姿可扩展性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。