用几何先验增强3D高斯点云的视角生成,减少幻觉并提升画质。
Enhancing Novel View Synthesis via Geometry Grounded Set Diffusion
- 基于集合的扩散模型融合几何先验与坐标对齐信息。
- 在多个数据集上实现感知保真度和结构相似性显著提升。
- 适合自动驾驶场景下的高质量多视角图像生成任务。
我们提出SetDiff,一种基于几何约束的多视角扩散框架,用于增强3D高斯点云生成的新视角渲染效果。该方法将显式的3D先验、像素对齐的坐标图以及姿态感知的Plucker射线嵌入整合到一个可处理任意数量参考视图与目标视图的集合式扩散模型中。该设计支持鲁棒的遮挡处理,在低信号条件下减少幻觉,并提升视觉内容恢复的光度保真度。统一的集合混合器在所有输入视图间执行全局令牌级注意力,支持可扩展的多相机增强,同时通过潜在空间监督与选择性解码保持计算效率。在EUVS、Para-Lane、nuScenes和DL3DV上的大量实验表明,该方法在严重外推条件下的感知保真度、结构相似性和鲁棒性方面均有显著提升。SetDiff为自动驾驶场景下真实且可靠的新型视角合成建立了新的扩散模型基准。
原文摘要 · Abstract (English)
We present SetDiff, a geometry-grounded multi-view diffusion framework that enhances novel-view renderings produced by 3D Gaussian Splatting. Our method integrates explicit 3D priors, pixel-aligned coordinate maps and pose-aware Plucker ray embeddings, into a set-based diffusion model capable of jointly processing variable numbers of reference and target views. This formulation enables robust occlusion handling, reduces hallucinations under low-signal conditions, and improves photometric fidelity in visual content restoration. A unified set mixer performs global token-level attention across all input views, supporting scalable multi-camera enhancement while maintaining computational efficiency through latent-space supervision and selective decoding. Extensive experiments on EUVS, Para-Lane, nuScenes, and DL3DV demonstrate significant gains in perceptual fidelity, structural similarity, and robustness under severe extrapolation. SetDiff establishes a state-of-the-art diffusion-based solution for realistic and reliable novel-view synthesis in autonomous driving scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。