MUSt3R让多视角图像直接生成统一坐标系下的3D重建,速度快且可扩展。
MUSt3R: Multi-view Network for Stereo 3D Reconstruction

- 将双视图模型升级为多视图对称网络,直接输出全局统一的3D结构
- 引入多层记忆机制,使千张图像重建在低复杂度下实现高速处理
- 适用于离线与在线场景,适合需要实时3D感知的SLAM和视觉里程计
DUSt3R提出了一种新范式,可在无相机标定或视角姿态先验的情况下,对任意图像集合进行密集且不受限的立体3D重建。然而,其通过处理图像对来回归局部3D结构,并需在全局坐标系中对齐,导致图像对数量随图像数平方增长,这对大规模图像集的鲁棒快速优化构成挑战。本文提出对DUSt3R的扩展:多视角立体3D重建网络MUSt3R,通过构建对称架构并直接在统一坐标系中预测所有视角的3D结构,解决了上述问题。其次,引入多层记忆机制,显著降低计算复杂度,使系统可在有限额外开销下实现数千张图像的高帧率3D点云生成。该框架支持离线与在线重建,适用于SfM和视觉SLAM场景,在未校准视觉里程计、相对相机位姿、尺度与焦距估计、3D重建及多视角深度估计等下游任务中表现达到当前最优水平。
原文摘要 · Abstract (English)
DUSt3R introduced a novel paradigm in geometric computer vision by proposing a model that can provide dense and unconstrained Stereo 3D Reconstruction of arbitrary image collections with no prior information about camera calibration nor viewpoint poses. Under the hood, however, DUSt3R processes image pairs, regressing local 3D reconstructions that need to be aligned in a global coordinate system. The number of pairs, growing quadratically, is an inherent limitation that becomes especially concerning for robust and fast optimization in the case of large image collections. In this paper, we propose an extension of DUSt3R from pairs to multiple views, that addresses all aforementioned concerns. Indeed, we propose a Multi-view Network for Stereo 3D Reconstruction, or MUSt3R, that modifies the DUSt3R architecture by making it symmetric and extending it to directly predict 3D structure for all views in a common coordinate frame. Second, we entail the model with a multi-layer memory mechanism which allows to reduce the computational complexity and to scale the reconstruction to large collections, inferring thousands of 3D pointmaps at high frame-rates with limited added complexity. The framework is designed to perform 3D reconstruction both offline and online, and hence can be seamlessly applied to SfM and visual SLAM scenarios showing state-of-the-art performance on various 3D downstream tasks, including uncalibrated Visual Odometry, relative camera pose, scale and focal estimation, 3D reconstruction and multi-view depth estimation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。