arXiv:2509.01873cs.CVcs.AI2025-09

融合几何先验的深度学习,提升3D视觉任务的精度与鲁棒性。

Doctoral Thesis: Geometric Deep Learning For Camera Pose Prediction, Registration, Depth Estimation, and 3D Reconstruction

  • 将深度、法向量等几何约束融入网络,增强模型对3D结构的理解。
  • 在真实场景中实现高保真3D重建,适用于文化遗产与VR/AR应用。
  • 解决无结构环境下的相机位姿估计与点云配准难题,适合3D视觉研究者。

现代深度学习的发展为3D地图构建、场景重建和虚拟现实开发带来了新机遇。尽管3D深度学习技术不断进步,直接对高维3D数据进行训练仍面临挑战,主要源于3D数据的高维度及标注数据稀缺。结构光(SfM)与同时定位与建图(SLAM)在结构化室内环境中表现稳健,但在非结构化环境中因特征模糊而性能下降,难以生成适用于渲染与语义分析的精细几何表示。当前局限要求发展融合传统几何方法与深度学习能力的3D表示技术,以构建具备几何感知能力的鲁棒模型。本论文针对3D视觉中的核心挑战,提出专用于相机位姿估计、点云配准、深度预测与高保真3D重建的几何深度学习方法。通过引入深度信息、表面法向量及等变性等几何先验,显著提升几何表征的准确性和鲁棒性。系统性研究了相机位姿估计、点云配准、深度估计与高质量3D重建,验证其在数字文化遗产保护与沉浸式VR/AR环境中的有效性。

原文摘要 · Abstract (English)

Modern deep learning developments create new opportunities for 3D mapping technology, scene reconstruction pipelines, and virtual reality development. Despite advances in 3D deep learning technology, direct training of deep learning models on 3D data faces challenges due to the high dimensionality inherent in 3D data and the scarcity of labeled datasets. Structure-from-motion (SfM) and Simultaneous Localization and Mapping (SLAM) exhibit robust performance when applied to structured indoor environments but often struggle with ambiguous features in unstructured environments. These techniques often struggle to generate detailed geometric representations effective for downstream tasks such as rendering and semantic analysis. Current limitations require the development of 3D representation methods that combine traditional geometric techniques with deep learning capabilities to generate robust geometry-aware deep learning models. The dissertation provides solutions to the fundamental challenges in 3D vision by developing geometric deep learning methods tailored for essential tasks such as camera pose estimation, point cloud registration, depth prediction, and 3D reconstruction. The integration of geometric priors or constraints, such as including depth information, surface normals, and equivariance into deep learning models, enhances both the accuracy and robustness of geometric representations. This study systematically investigates key components of 3D vision, including camera pose estimation, point cloud registration, depth estimation, and high-fidelity 3D reconstruction, demonstrating their effectiveness across real-world applications such as digital cultural heritage preservation and immersive VR/AR environments.

3D重建几何深度学习相机位姿点云配准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。