用多模态先验引导采样,提升稀疏视角新视图合成的细节重建质量。
Multimodal-Prior-Guided Importance Sampling for Hierarchical Gaussian Splatting in Sparse-View Novel View Synthesis
- 融合光照、语义和几何先验,动态判断何处注入精细高斯点。
- 在DTU数据集上达到最高+0.3 dB PSNR,优于现有方法。
- 适合需要高保真3D重建的视觉建模与自动驾驶场景应用。
我们提出以多模态先验引导的重要性采样作为稀疏视角新视图合成中分层3D高斯溅射的核心机制。该采样器融合光照渲染残差、语义先验和几何先验,生成稳健的局部可恢复性估计,直接指导精细高斯点的注入位置。基于此采样核心,我们的框架包含:(1) 从粗到细的高斯表示,通过稳定粗粒度层编码全局形状,并仅在多模态指标指示可恢复细节处添加精细原型;(2) 几何感知的采样与保留策略,聚焦于几何关键与复杂区域的优化,同时保护欠约束区域中新添加的原型免于过早剪枝。通过优先考虑一致的多模态证据而非仅依赖原始残差,本方法缓解了纹理引起的过拟合误差,并抑制了姿态/外观不一致带来的噪声。在多种稀疏视角基准上的实验表明,本方法实现了当前最优重建效果,其中在DTU数据集上达到最高+0.3 dB PSNR。
原文摘要 · Abstract (English)
We present multimodal-prior-guided importance sampling as the central mechanism for hierarchical 3D Gaussian Splatting (3DGS) in sparse-view novel view synthesis. Our sampler fuses complementary cues { -- } photometric rendering residuals, semantic priors, and geometric priors { -- } to produce a robust, local recoverability estimate that directly drives where to inject fine Gaussians. Built around this sampling core, our framework comprises (1) a coarse-to-fine Gaussian representation that encodes global shape with a stable coarse layer and selectively adds fine primitives where the multimodal metric indicates recoverable detail; and (2) a geometric-aware sampling and retention policy that concentrates refinement on geometrically critical and complex regions while protecting newly added primitives in underconstrained areas from premature pruning. By prioritizing regions supported by consistent multimodal evidence rather than raw residuals alone, our method alleviates overfitting texture-induced errors and suppresses noise from pose/appearance inconsistencies. Experiments on diverse sparse-view benchmarks demonstrate state-of-the-art reconstructions, with up to +0.3 dB PSNR on DTU.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。