用视频扩散模型补全少视角图像,实现高保真3D场景重建
SpatialCrafter: Unleashing the Imagination of Video Diffusion Models for Scene Reconstruction from Limited Observations
- 利用视频扩散模型生成合理额外视角,缓解单/稀疏视图下的重建歧义
- 通过可训练相机编码器与对极注意力机制,实现精准相机控制与3D一致性
- 支持长序列处理,融合单目深度与语义特征,直接回归3D高斯实例
新视角合成(NVS)在计算机视觉与图形学中推动沉浸式体验发展。现有方法虽有进展,但仍依赖密集多视角观测,限制了实际应用。本文挑战从稀疏或单视图输入重建逼真3D场景的问题。提出SpatialCrafter框架,利用视频扩散模型中的丰富先验知识生成合理的附加观测,降低重建不确定性。通过可训练的相机编码器与显式几何约束的对极注意力机制,实现精确相机控制和3D一致性,并采用统一尺度估计策略解决不同数据集间的尺度差异。此外,结合单目深度先验与视频潜在空间中的语义特征,直接回归3D高斯原型,并通过混合网络结构高效处理长序列特征。大量实验表明,该方法显著提升稀疏视图重建质量,恢复出真实感的3D场景外观。
原文摘要 · Abstract (English)
Novel view synthesis (NVS) boosts immersive experiences in computer vision and graphics. Existing techniques, though progressed, rely on dense multi-view observations, restricting their application. This work takes on the challenge of reconstructing photorealistic 3D scenes from sparse or single-view inputs. We introduce SpatialCrafter, a framework that leverages the rich knowledge in video diffusion models to generate plausible additional observations, thereby alleviating reconstruction ambiguity. Through a trainable camera encoder and an epipolar attention mechanism for explicit geometric constraints, we achieve precise camera control and 3D consistency, further reinforced by a unified scale estimation strategy to handle scale discrepancies across datasets. Furthermore, by integrating monocular depth priors with semantic features in the video latent space, our framework directly regresses 3D Gaussian primitives and efficiently processes long-sequence features using a hybrid network structure. Extensive experiments show our method enhances sparse view reconstruction and restores the realistic appearance of 3D scenes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。