用少量杂乱真实图像实现高质量3D重建,解决遮挡与外观变化问题。
Difix3D-W: Distractor-Free Few-Shot 3D Gaussian Splatting in the Wild
- 引入参考图与瞬态掩码,一步扩散模型优化视角渲染
- 通过稀疏区域增强策略提升高斯点密度,改善视角缺失问题
- 结合LoRA与正则化保持多视角一致性,适合真实场景应用
我们提出Difix3D-W,一种面向无约束真实场景的3D稀疏视角合成框架,可处理包含干扰物、遮挡和外观变化的情况。不同于以往仅在受限图像集上进行新视角合成,或依赖真实世界中密集图像集合的方法,本方法使用稀疏非约束图像即可生成高质量3D渲染结果。为此,我们设计了一种基于参考图的视角精炼机制,采用改进的一步扩散模型并结合瞬态掩码与参考图像,有效减少渲染伪影,增强高斯场中的3D表示。此外,通过稀疏感知的高斯复制策略放大稀疏区域的高斯点,缓解相机视角不足的问题。最后,利用LoRA与正则化手段维持3D多视角一致性。大量实验表明,该方法持续优于现有方法,为无需繁重数据采集的真实场景3D重建提供了新路径。
原文摘要 · Abstract (English)
We propose Difix3D-W, a 3D novel sparse-view synthesis framework for unconstrained real-world scenarios that contain distractors, occlusion, and appearance variation. Unlike existing methods that primarily perform novel-view synthesis from a sparse set of constrained images without transient elements or leverage unconstrained dense image collections in real-world scenarios, our method utilize sparse unconstrained images, showing high-quality 3D rendering results. To do this, we introduce reference-guided view refinement with a redesigned one-step diffusion model using a transient mask and a reference image to mitigate artifacts in rendered views, enhancing the 3D representation in the Gaussian field. Furthermore, we address sparse regions in the Gaussian field leveraging sparsity-aware Gaussian replication strategy to amplify Gaussians in the sparse regions and alleviate deficient camera viewpoint issues. Finally, we utilize LoRA and regularization to maintain 3D multi-view consistency. Extensive experiments demonstrate that our method consistently outperforms existing methods. This advancement paves the way for realizing real-world scenarios without labor-intensive data acquisition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。