用3D视觉基础模型实现无需约束的快速结构光恢复
MASt3R-SfM: a Fully-Integrated Solution for Unconstrained Structure-from-Motion

- 基于3D视觉基础模型生成局部重建与匹配
- 全局对齐仅需低内存,复杂度从二次降为线性
- 适合任意图像集合,小中规模表现更优
结构光恢复(SfM)旨在从一组图像中联合恢复相机位姿和场景三维几何,尽管经过数十年发展仍面临诸多挑战。传统方法依赖复杂的最小解算器流水线,易传播误差,且在图像重叠不足或运动过小时失效。近期方法虽尝试重构该范式,但实证表明其未能解决核心问题。本文提出基于最新发布的3D视觉基础模型,该模型可鲁棒生成局部3D重建与精准匹配。我们引入一种低内存方法,将这些局部重建准确对齐至全局坐标系,并进一步证明该基础模型可作为高效图像检索器,无额外开销,使整体复杂度从二次降至线性。所提SfM流水线简单、可扩展、快速且真正无约束,能处理任意图像集合(有序或无序)。多基准测试表明,本方法在不同设置下性能稳定,尤其在小中规模场景中显著优于现有方法。
原文摘要 · Abstract (English)
Structure-from-Motion (SfM), a task aiming at jointly recovering camera poses and 3D geometry of a scene given a set of images, remains a hard problem with still many open challenges despite decades of significant progress. The traditional solution for SfM consists of a complex pipeline of minimal solvers which tends to propagate errors and fails when images do not sufficiently overlap, have too little motion, etc. Recent methods have attempted to revisit this paradigm, but we empirically show that they fall short of fixing these core issues. In this paper, we propose instead to build upon a recently released foundation model for 3D vision that can robustly produce local 3D reconstructions and accurate matches. We introduce a low-memory approach to accurately align these local reconstructions in a global coordinate system. We further show that such foundation models can serve as efficient image retrievers without any overhead, reducing the overall complexity from quadratic to linear. Overall, our novel SfM pipeline is simple, scalable, fast and truly unconstrained, i.e. it can handle any collection of images, ordered or not. Extensive experiments on multiple benchmarks show that our method provides steady performance across diverse settings, especially outperforming existing methods in small- and medium-scale settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。