对比传统与学习型三维重建方法,发现后者更快更鲁棒,但精度略低。
A Comparison of Multi-View Stereo Methods for Photogrammetric 3D Reconstruction: From Traditional to Learning-Based Approaches

- 对比COLMAP与多种学习型方法在真实场景下的表现
- 学习型方法速度更快,传统方法几何更一致
- 适合需要快速重建的无人机或大规模场景应用
摄影测量三维重建长期依赖传统的结构光(SfM)与多视图立体(MVS)方法,虽精度高但存在速度慢、可扩展性差的问题。近期学习型MVS方法兴起,旨在实现更快更高效的重建。本文对比了代表性传统MVS流程(COLMAP)与前沿学习型方法,包括基于几何引导的方法(MVSNet、PatchmatchNet、MVSAnywhere、MVSFormer++)和端到端框架(Stereo4D、FoundationStereo、DUSt3R、MASt3R、Fast3R、VGGT)。在两个不同航拍场景下进行实验:第一使用MARS-LVIG数据集,以激光雷达点云为真值;第二使用Pix4D官网公开场景,真值由Pix4Dmapper生成。评估了精度、覆盖率与运行时间。结果表明,尽管COLMAP能提供可靠且几何一致的重建,但耗时更长;当传统方法在图像配准中失败时,学习型方法表现出更强特征匹配能力与更高鲁棒性。基于几何引导的方法通常需精细数据准备,并依赖由COLMAP生成的相机位姿或深度先验。端到端方法如DUSt3R与VGGT在保持较优精度与合理覆盖率的同时,显著提升重建速度,但在挑战性场景中仍存在较大重建残差。
原文摘要 · Abstract (English)
Photogrammetric 3D reconstruction has long relied on traditional Structure-from-Motion (SfM) and Multi-View Stereo (MVS) methods, which provide high accuracy but face challenges in speed and scalability. Recently, learning-based MVS methods have emerged, aiming for faster and more efficient reconstruction. This work presents a comparative evaluation between a representative traditional MVS pipeline (COLMAP) and state-of-the-art learning-based approaches, including geometry-guided methods (MVSNet, PatchmatchNet, MVSAnywhere, MVSFormer++) and end-to-end frameworks (Stereo4D, FoundationStereo, DUSt3R, MASt3R, Fast3R, VGGT). Two experiments were conducted on different aerial scenarios. The first experiment used the MARS-LVIG dataset, where ground-truth 3D reconstruction was provided by LiDAR point clouds. The second experiment used a public scene from the Pix4D official website, with ground truth generated by Pix4Dmapper. We evaluated accuracy, coverage, and runtime across all methods. Experimental results show that although COLMAP can provide reliable and geometrically consistent reconstruction results, it requires more computation time. In cases where traditional methods fail in image registration, learning-based approaches exhibit stronger feature-matching capability and greater robustness. Geometry-guided methods usually require careful dataset preparation and often depend on camera pose or depth priors generated by COLMAP. End-to-end methods such as DUSt3R and VGGT achieve competitive accuracy and reasonable coverage while offering substantially faster reconstruction. However, they exhibit relatively large residuals in 3D reconstruction, particularly in challenging scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。