融合点云投射与扩散模型,实现单图高保真新视角生成
High-Fidelity Novel View Synthesis via Splatting-Guided Diffusion
- 用点云投射引导扩散模型,精准控制视角与几何
- 引入纹理桥模块,有效避免纹理幻觉,提升细节真实感
- 无需微调即可跨任务零样本应用,适合多场景生成
尽管新视角合成(NVS)近期取得进展,但从单图或稀疏观测生成高保真视图仍是重大挑战。现有基于点云投射的方法常因投射误差导致几何失真,而扩散模型虽能利用丰富3D先验改善几何,却易产生纹理幻觉。本文提出SplatDiff,一种像素级点云投射引导的视频扩散模型,可从单图生成高保真新视角。我们设计了对齐合成策略以精确控制目标视角,并实现几何一致的视图合成;为缓解纹理幻觉,提出纹理桥模块,通过自适应特征融合实现高保真纹理生成。SplatDiff结合点云投射与扩散模型优势,在单视图NVS上达到顶尖性能。此外,无需额外训练,其在稀疏视图NVS和立体视频转换等多样化任务中展现出卓越零样本能力。
原文摘要 · Abstract (English)
Despite recent advances in Novel View Synthesis (NVS), generating high-fidelity views from single or sparse observations remains a significant challenge. Existing splatting-based approaches often produce distorted geometry due to splatting errors. While diffusion-based methods leverage rich 3D priors to achieve improved geometry, they often suffer from texture hallucination. In this paper, we introduce SplatDiff, a pixel-splatting-guided video diffusion model designed to synthesize high-fidelity novel views from a single image. Specifically, we propose an aligned synthesis strategy for precise control of target viewpoints and geometry-consistent view synthesis. To mitigate texture hallucination, we design a texture bridge module that enables high-fidelity texture generation through adaptive feature fusion. In this manner, SplatDiff leverages the strengths of splatting and diffusion to generate novel views with consistent geometry and high-fidelity details. Extensive experiments verify the state-of-the-art performance of SplatDiff in single-view NVS. Additionally, without extra training, SplatDiff shows remarkable zero-shot performance across diverse tasks, including sparse-view NVS and stereo video conversion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。