用视频扩散模型补全稀疏视角下的3D场景,无需训练即可还原完整结构。
VidSplat: Gaussian Splatting Reconstruction with Geometry-Guided Video Diffusion Priors

- 基于几何引导的分阶段去噪,实现3D一致的生成。
- 通过迭代采样相机轨迹,合成新视角并逐步优化重建结果。
- 仅需单张图或稀疏视角即可完成高质量3D重建,适合低数据场景。
高斯点阵在多视角表面重建中取得显著进展,但在仅有少数视角时表现明显下降。尽管近期工作通过增强多视角一致性生成合理表面,仍难以推断未被观测、遮挡或约束薄弱的区域。为此,我们提出VidSplat——一种无需训练的生成式重建框架,利用强大的视频扩散先验,迭代合成补充视角以弥补输入覆盖不足,从而从稀疏输入中恢复完整3D场景。具体而言,为有效融合生成与重建,我们解决两大挑战:其一,针对3D一致性生成,设计一种无需训练的分阶段去噪策略,通过渲染的RGB与掩码图像自适应引导去噪方向朝向底层几何;其二,为提升重建效果,提出迭代机制,采样相机轨迹,探索未观测区域,合成新视角,并通过置信度加权精炼进行补充训练。VidSplat在稀疏输入下表现稳健,甚至可处理单张图像。广泛基准测试表明,其在稀疏视角场景重建上优于现有方法。
原文摘要 · Abstract (English)
Gaussian Splatting has achieved remarkable progress in multi-view surface reconstruction, yet it exhibits notable degradation when only few views are available. Although recent efforts alleviate this issue by enhancing multi-view consistency to produce plausible surfaces, they struggle to infer unseen, occluded, or weakly constrained regions beyond the input coverage. To address this limitation, we present VidSplat, a training-free generative reconstruction framework that leverages powerful video diffusion priors to iteratively synthesize novel views that compensate for missing input coverage, and thereby recover complete 3D scenes from sparse inputs. Specifically, we tackle two key challenges that enable the effective integration of generation and reconstruction. First, for 3D consistent generation, we elaborate a training-free, stage-wise denoising strategy that adaptively guides the denoising direction toward the underlying geometry using the rendered RGB and mask images. Second, to enhance the reconstruction, we develop an iterative mechanism that samples camera trajectories, explores unobserved regions, synthesizes novel views, and supplements training through confidence weighted refinement. VidSplat performs robustly to sparse input and even a single image. Extensive experiments on widely used benchmarks demonstrate our superior performance in sparse-view scene reconstruction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。