单目视频实时3D重建,速度近10帧每秒
MAGiSt3R: Multi-Agent Feed-forward 3D Reconstruction from Monocular RGB Videos

- 多智能体并行处理,分块回归点云图
- 融合本地与全局信息,重建精度优于现有方法
- 适合需要高速高精度3D重建的场景
本文提出MAGiSt3R,一种基于单目RGB视频的多智能体3D重建框架,可在接近10 FPS的速度下完成重建与相机跟踪。该框架采用3R系列的前馈模型处理视频并回归局部点云图,并通过名为MAGMA的融合模型,在智能体内与跨智能体层面整合局部地图,生成最终全局点云图。此外,MAGiSt3R引入位姿图优化,有效缓解前馈流程中的累积相机漂移问题。我们在合成与真实世界数据集上进行了评估,结果表明其在重建与相机跟踪精度上均优于当前最优方法。
原文摘要 · Abstract (English)
This paper presents MAGiSt3R, a multi-agent 3D reconstruction framework performing reconstruction and camera tracking for monocular RGB videos at almost 10 FPS. MAGiSt3R relies on a feed-forward model from the 3R family to process RGB videos and regress local point maps, and on a merging model, MAGMA, that combines local maps at both intra-agent and inter-agent levels to obtain the final global point map. Furthermore, MAGiSt3R performs pose graph optimization to mitigate cumulative camera drift occurring along the feed-forward pipeline. We evaluate MAGiSt3R on both synthetic and real-world datasets, demonstrating its superior reconstruction and camera tracking accuracy compared to state-of-the-art approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。