用生成先验和异常检测,提升稀疏视角下的3D重建与姿态估计精度。
Sparse-view Pose Estimation and Reconstruction via Analysis by Generative Synthesis
- 结合新视角合成与光度损失,引入生成先验优化3D结构
- 通过离散搜索+连续优化,有效修正姿态估计中的异常值
- 适用于真实与合成数据,显著提升现有系统性能
从多视角图像推断3D结构通常需协同解决3D重建与相机姿态估计问题。传统分析合成框架将此视为联合优化问题,近期方法利用神经场等表达并基于梯度下降优化初始姿态。然而,在稀疏视角下,观测证据不足导致3D信息不完整,姿态误差难以纠正且会进一步恶化重建结果。为此,本文提出SparseAGS,通过:a) 结合新视角合成的生成先验与光度目标以提升3D质量;b) 显式建模异常值,采用离散搜索与连续优化相结合策略进行修正。我们在真实世界与合成数据集上验证了该框架,结合多种现成姿态估计系统作为初始化。结果表明,该方法显著提升了基线系统的姿态精度,并生成了优于当前多视图重建基线的高质量3D重建结果。
原文摘要 · Abstract (English)
Inferring the 3D structure underlying a set of multi-view images typically requires solving two co-dependent tasks -- accurate 3D reconstruction requires precise camera poses, and predicting camera poses relies on (implicitly or explicitly) modeling the underlying 3D. The classical framework of analysis by synthesis casts this inference as a joint optimization seeking to explain the observed pixels, and recent instantiations learn expressive 3D representations (e.g., Neural Fields) with gradient-descent-based pose refinement of initial pose estimates. However, given a sparse set of observed views, the observations may not provide sufficient direct evidence to obtain complete and accurate 3D. Moreover, large errors in pose estimation may not be easily corrected and can further degrade the inferred 3D. To allow robust 3D reconstruction and pose estimation in this challenging setup, we propose SparseAGS, a method that adapts this analysis-by-synthesis approach by: a) including novel-view-synthesis-based generative priors in conjunction with photometric objectives to improve the quality of the inferred 3D, and b) explicitly reasoning about outliers and using a discrete search with a continuous optimization-based strategy to correct them. We validate our framework across real-world and synthetic datasets in combination with several off-the-shelf pose estimation systems as initialization. We find that it significantly improves the base systems' pose accuracy while yielding high-quality 3D reconstructions that outperform the results from current multi-view reconstruction baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。