用一步扩散模型提升稀疏图像新视角生成的细节质量
One-Shot Refiner: Boosting Feed-forward Novel View Synthesis via One-Step Diffusion
- 双域感知模块让高分辨率输入不再受ViT限制
- 扩散网络有效保留高频细节,避免跨视图结构不一致
- 联合训练使几何与细节修复协同优化,适合高质量3D重建
我们提出一种新型框架,用于从稀疏图像中实现高保真新视角合成(NVS),解决了基于视觉变换器(ViT)骨干网络的前馈式3D高斯点云(3DGS)方法的关键局限。尽管ViT管道具备强几何先验,但常受限于低分辨率输入,因计算开销大。此外,现有生成增强方法多为3D无关,导致视图间结构不一致,尤其在未见区域更明显。为此,我们设计了双域细节感知模块,使高分辨率图像处理不受限于ViT骨干,同时为高斯点赋予额外特征以存储高频细节。我们开发了特征引导的扩散网络,在恢复过程中保持高频细节。引入统一训练策略,实现ViT几何骨干与扩散增强模块的联合优化。实验表明,该方法在多个数据集上均能维持优越生成质量。
原文摘要 · Abstract (English)
We present a novel framework for high-fidelity novel view synthesis (NVS) from sparse images, addressing key limitations in recent feed-forward 3D Gaussian Splatting (3DGS) methods built on Vision Transformer (ViT) backbones. While ViT-based pipelines offer strong geometric priors, they are often constrained by low-resolution inputs due to computational costs. Moreover, existing generative enhancement methods tend to be 3D-agnostic, resulting in inconsistent structures across views, especially in unseen regions. To overcome these challenges, we design a Dual-Domain Detail Perception Module, which enables handling high-resolution images without being limited by the ViT backbone, and endows Gaussians with additional features to store high-frequency details. We develop a feature-guided diffusion network, which can preserve high-frequency details during the restoration process. We introduce a unified training strategy that enables joint optimization of the ViT-based geometric backbone and the diffusion-based refinement module. Experiments demonstrate that our method can maintain superior generation quality across multiple datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。