用稀疏航拍图测试三款新3D重建模型,发现它们在低重叠场景下表现优异。
An Evaluation of DUSt3R/MASt3R/VGGT 3D Reconstruction on Photogrammetric Aerial Blocks
- 用预训练模型直接处理极稀疏航拍图像,跳过传统流程
- 少于10张图也能生成完整点云,完整度比COLMAP高50%
- 适合低分辨率、纹理少或重叠差的复杂航拍场景
当前最先进的3D视觉算法在处理稀疏无序图像集方面持续进步。近期推出的三大基础模型——密集无约束立体3D重建(DUSt3R)、匹配与立体3D重建(MASt3R)和视觉几何接地变压器(VGGT),因其能处理极低图像重叠而受到关注。评估这些模型在典型航拍图像上的表现至关重要,因它们可能应对极低重叠、立体遮挡和无纹理区域。对于冗余数据集,可使用极度稀疏的图像集加速3D重建。尽管已在多个计算机视觉基准上测试,其在摄影测量航拍块中的潜力仍不明。本文对UseGeo数据集的航拍块,全面评估了预训练的DUSt3R/MASt3R/VGGT模型在位姿估计与密集3D重建上的表现。结果表明,这些方法可从极少图像(少于10张,最高518像素分辨率)中准确重建密集点云,完整性最高较COLMAP提升50%。VGGT还表现出更高计算效率、可扩展性及更可靠的相机位姿估计。但所有模型在高分辨率图像和大规模数据集上均有局限,位姿可靠性随图像数量和几何复杂度增加而下降。研究显示,基于Transformer的方法尚不能完全替代传统SfM和MVS,但在挑战性、低分辨率和稀疏场景中具有互补潜力。
原文摘要 · Abstract (English)
State-of-the-art 3D computer vision algorithms continue to advance in handling sparse, unordered image sets. Recently developed foundational models for 3D reconstruction, such as Dense and Unconstrained Stereo 3D Reconstruction (DUSt3R), Matching and Stereo 3D Reconstruction (MASt3R), and Visual Geometry Grounded Transformer (VGGT), have attracted attention due to their ability to handle very sparse image overlaps. Evaluating DUSt3R/MASt3R/VGGT on typical aerial images matters, as these models may handle extremely low image overlaps, stereo occlusions, and textureless regions. For redundant collections, they can accelerate 3D reconstruction by using extremely sparsified image sets. Despite tests on various computer vision benchmarks, their potential on photogrammetric aerial blocks remains unexplored. This paper conducts a comprehensive evaluation of the pre-trained DUSt3R/MASt3R/VGGT models on the aerial blocks of the UseGeo dataset for pose estimation and dense 3D reconstruction. Results show these methods can accurately reconstruct dense point clouds from very sparse image sets (fewer than 10 images, up to 518 pixels resolution), with completeness gains up to +50% over COLMAP. VGGT also demonstrates higher computational efficiency, scalability, and more reliable camera pose estimation. However, all exhibit limitations with high-resolution images and large sets, as pose reliability declines with more images and geometric complexity. These findings suggest transformer-based methods cannot fully replace traditional SfM and MVS, but offer promise as complementary approaches, especially in challenging, low-resolution, and sparse scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。