用可学习的注意力机制替代传统优化,实现高效3D重建。
Light3R-SfM: Towards Feed-forward Structure-from-Motion
- 用可学习的注意力模块替代全局优化,实现前馈式建模。
- 通过检索得分引导的最短路径树,内存占用大幅降低。
- 适合实时性要求高的真实场景3D重建任务。
我们提出 Light3R-SfM,一种面向无约束图像集合的前馈、端到端可学习的高效大规模结构光流(SfM)框架。与依赖昂贵匹配和全局优化的传统SfM方法不同,Light3R-SfM通过新颖的潜在全局对齐模块,将传统全局优化替换为可学习的注意力机制,有效捕捉多视角约束,实现鲁棒且精确的相机位姿估计。该方法通过检索得分引导的最短路径树构建稀疏场景图,在内存使用和计算开销上相比朴素方法显著降低。大量实验表明,Light3R-SfM在保持竞争力精度的同时,显著减少运行时间,适用于具有实时性约束的真实世界3D重建任务。本工作开创了数据驱动的前馈式SfM范式,为野外环境下可扩展、准确且高效的3D重建铺平道路。
原文摘要 · Abstract (English)
We present Light3R-SfM, a feed-forward, end-to-end learnable framework for efficient large-scale Structure-from-Motion (SfM) from unconstrained image collections. Unlike existing SfM solutions that rely on costly matching and global optimization to achieve accurate 3D reconstructions, Light3R-SfM addresses this limitation through a novel latent global alignment module. This module replaces traditional global optimization with a learnable attention mechanism, effectively capturing multi-view constraints across images for robust and precise camera pose estimation. Light3R-SfM constructs a sparse scene graph via retrieval-score-guided shortest path tree to dramatically reduce memory usage and computational overhead compared to the naive approach. Extensive experiments demonstrate that Light3R-SfM achieves competitive accuracy while significantly reducing runtime, making it ideal for 3D reconstruction tasks in real-world applications with a runtime constraint. This work pioneers a data-driven, feed-forward SfM approach, paving the way toward scalable, accurate, and efficient 3D reconstruction in the wild.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。