用视频生成模型指导稀疏视角3D高斯点云,提升重建质量。
RealisticDreamer: Guidance Score Distillation for Few-shot Gaussian Splatting
- 从预训练视频扩散模型中提取多视角一致性先验作为指导信号
- 在多个邻近视角上监督渲染图像,减少稀疏视图下的过拟合
- 结合深度图与语义特征修正噪声预测,适配真实相机位姿
3D高斯点云(3DGS)因其高质量实时渲染能力,在3D场景表示中受到广泛关注。然而,当输入训练视角稀疏时,3DGS易发生过拟合,主要因缺乏中间视角的监督。受视频扩散模型(VDM)成功启发,本文提出引导得分蒸馏(GSD)框架,从预训练VDM中提取丰富的多视角一致性先验。基于得分蒸馏采样(SDS)的洞察,GSD通过多个邻近视角的渲染图像进行监督,引导高斯点云表示向VDM的生成方向演化。但生成方向常涉及物体运动和随机相机轨迹,直接优化难以对齐。为此,本文引入统一引导形式,修正VDM的噪声预测结果,结合真实深度图的深度映射引导与基于语义图像特征的引导,确保得分更新方向与正确相机姿态和精确几何一致。实验表明,本方法在多个数据集上均优于现有方法。
原文摘要 · Abstract (English)
3D Gaussian Splatting (3DGS) has recently gained great attention in the 3D scene representation for its high-quality real-time rendering capabilities. However, when the input comprises sparse training views, 3DGS is prone to overfitting, primarily due to the lack of intermediate-view supervision. Inspired by the recent success of Video Diffusion Models (VDM), we propose a framework called Guidance Score Distillation (GSD) to extract the rich multi-view consistency priors from pretrained VDMs. Building on the insights from Score Distillation Sampling (SDS), GSD supervises rendered images from multiple neighboring views, guiding the Gaussian splatting representation towards the generative direction of VDM. However, the generative direction often involves object motion and random camera trajectories, making it challenging for direct supervision in the optimization process. To address this problem, we introduce an unified guidance form to correct the noise prediction result of VDM. Specifically, we incorporate both a depth warp guidance based on real depth maps and a guidance based on semantic image features, ensuring that the score update direction from VDM aligns with the correct camera pose and accurate geometry. Experimental results show that our method outperforms existing approaches across multiple datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。