解决野外稀疏多姿态图像的3D重建难题,实现逼真新视角合成。
MS-GS: Multi-Appearance Sparse-View 3D Gaussian Splatting in the Wild
- 基于单目深度先验和SfM点锚定,构建局部语义区域提升对齐精度。
- 在虚拟视图上施加像素与特征级几何引导,减少过拟合并增强3D一致性。
- 适用于光照/季节变化大的真实场景,适合野外图像重建任务。
野外照片集常因图像数量有限且呈现多姿态(如不同时间段或季节),给场景重建与新视角合成带来挑战。尽管近期神经辐射场(NeRF)和3D高斯泼溅(3DGS)的改进版本有所提升,但仍易出现过度平滑与过拟合问题。本文提出MS-GS,一种面向稀疏视图下多姿态场景的新型3DGS框架。为弥补初始稀疏性带来的不足,方法利用单目深度估计提取的几何先验。核心在于通过结构光流(SfM)点锚定算法识别并利用局部语义区域,实现可靠对齐与几何线索捕捉。进一步,设计一系列在虚拟视图上的像素与特征层级几何引导监督步骤,强化3D一致性、抑制过拟合。同时引入新数据集与真实场景实验设置,建立更具挑战性的基准。结果表明,MS-GS在多种稀疏视图与多姿态条件下均能生成逼真渲染效果,在多个数据集上显著优于现有方法。
原文摘要 · Abstract (English)
In-the-wild photo collections often contain limited volumes of imagery and exhibit multiple appearances, e.g., taken at different times of day or seasons, posing significant challenges to scene reconstruction and novel view synthesis. Although recent adaptations of Neural Radiance Field (NeRF) and 3D Gaussian Splatting (3DGS) have improved in these areas, they tend to oversmooth and are prone to overfitting. In this paper, we present MS-GS, a novel framework designed with Multi-appearance capabilities in Sparse-view scenarios using 3DGS. To address the lack of support due to sparse initializations, our approach is built on the geometric priors elicited from monocular depth estimations. The key lies in extracting and utilizing local semantic regions with a Structure-from-Motion (SfM) points anchored algorithm for reliable alignment and geometry cues. Then, to introduce multi-view constraints, we propose a series of geometry-guided supervision steps at virtual views in pixel and feature levels to encourage 3D consistency and reduce overfitting. We also introduce a dataset and an in-the-wild experiment setting to set up more realistic benchmarks. We demonstrate that MS-GS achieves photorealistic renderings under various challenging sparse-view and multi-appearance conditions, and outperforms existing approaches significantly across different datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。