用3D基础模型隐式表示生成新视角深度图,无需训练。
Zero-Shot Novel Depth Synthesis Using 3D Foundation Models Scene Representations

- 利用3D基础模型内部表征进行潜在扩散生成新视图点云。
- 在多个数据集上生成真实感深度图,零样本泛化能力强。
- 适合研究3D生成、新视角合成与基础模型应用的开发者。
近期的3D基础模型(如VGGT)通过前馈变换器预测出丰富的统一场景表征,在多项3D视觉任务中表现优异。本文探究利用其内部表征从新视角推断3D结构的可能性。我们假设,为实现3D重建任务,这些模型需学习包含大量通用3D场景知识的表征。实验表明,可从3DFM内部表征中解码隐藏表面。为此,我们提出Z3D方法,通过在3DFM表征上执行潜在扩散,估计未见视角的点图。结果表明,Z3D可在多个数据集上生成真实感深度图,实现零样本新视角深度合成。
原文摘要 · Abstract (English)
3D Foundation Models (3DFMs) such as VGGT have recently pushed the boundaries of 3D vision by predicting rich unified representations with feed-foward transformers. The scene representations learned by these models enable strong performance on multiple 3D vision tasks. In this paper, we investigate using their internal representations to infer 3D in the scene from new views. Our hypothesis is that in order to solve the task of 3D reconstruction, these models need to learn a representation that includes a large amount of general knowledge about 3D scenes. After showing that it is possible to decode hidden surfaces from internal 3DFM representations, we propose a method, Z3D, that estimates pointmaps in unseen views by doing latent diffusion on 3DFM representation. We show that Z3D can predict realistic depth maps for new views across multiple datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。