arXiv:2607.15211cs.CV2026-07被引 1

单目视频实时3D重建,速度近10帧每秒

MAGiSt3R: Multi-Agent Feed-forward 3D Reconstruction from Monocular RGB Videos

论文配图:MAGiSt3R: Multi-Agent Feed-forward 3D Reconstruction from Monocular RGB Videos
图 1 · 摘自论文原文
  • 多智能体并行处理,分块回归点云图
  • 融合本地与全局信息,重建精度优于现有方法
  • 适合需要高速高精度3D重建的场景

本文提出MAGiSt3R,一种基于单目RGB视频的多智能体3D重建框架,可在接近10 FPS的速度下完成重建与相机跟踪。该框架采用3R系列的前馈模型处理视频并回归局部点云图,并通过名为MAGMA的融合模型,在智能体内与跨智能体层面整合局部地图,生成最终全局点云图。此外,MAGiSt3R引入位姿图优化,有效缓解前馈流程中的累积相机漂移问题。我们在合成与真实世界数据集上进行了评估,结果表明其在重建与相机跟踪精度上均优于当前最优方法。

原文摘要 · Abstract (English)

This paper presents MAGiSt3R, a multi-agent 3D reconstruction framework performing reconstruction and camera tracking for monocular RGB videos at almost 10 FPS. MAGiSt3R relies on a feed-forward model from the 3R family to process RGB videos and regress local point maps, and on a merging model, MAGMA, that combines local maps at both intra-agent and inter-agent levels to obtain the final global point map. Furthermore, MAGiSt3R performs pose graph optimization to mitigate cumulative camera drift occurring along the feed-forward pipeline. We evaluate MAGiSt3R on both synthetic and real-world datasets, demonstrating its superior reconstruction and camera tracking accuracy compared to state-of-the-art approaches.

3D重建单目视频多智能体实时

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。