arXiv:2603.16844cs.CV2026-03

用多视图大模型提升单目重建精度,实现更稳定的实时三维重建。

M^3: Dense Matching Meets Multi-View Foundation Models for Monocular Gaussian Splatting SLAM

  • 在多视图大模型中加入匹配头,生成精细稠密对应关系。
  • 相比VGGT-SLAM 2.0,ATE RMSE降低64.3%,ScanNet++上PSNR提升2.11 dB。
  • 适合需要高精度单目重建的动态环境应用,如机器人导航与AR。

从未标定的单目视频流进行实时重建仍具挑战,需在动态环境中实现高精度位姿估计与高效在线优化。尽管将3D基础模型与SLAM框架结合是可行方向,但关键瓶颈在于:多数多视图基础模型采用前馈方式估计位姿,生成的像素级对应关系精度不足,难以支撑严格的几何优化。为此,本文提出M^3,通过添加专用匹配头增强多视图基础模型,实现细粒度稠密对应关系,并集成至鲁棒的单目高斯泼溅SLAM系统。M^3还通过动态区域抑制与跨推理内在对齐提升跟踪稳定性。在多样化的室内外基准测试中,其位姿估计与场景重建均达到当前最优水平。特别地,在ScanNet++数据集上,相比VGGT-SLAM 2.0,ATE RMSE降低64.3%;在PSNR上优于ARTDECO 2.11 dB。

原文摘要 · Abstract (English)

Streaming reconstruction from uncalibrated monocular video remains challenging, as it requires both high-precision pose estimation and computationally efficient online refinement in dynamic environments. While coupling 3D foundation models with SLAM frameworks is a promising paradigm, a critical bottleneck persists: most multi-view foundation models estimate poses in a feed-forward manner, yielding pixel-level correspondences that lack the requisite precision for rigorous geometric optimization. To address this, we present M^3, which augments the Multi-view foundation model with a dedicated Matching head to facilitate fine-grained dense correspondences and integrates it into a robust Monocular Gaussian Splatting SLAM. M^3 further enhances tracking stability by incorporating dynamic area suppression and cross-inference intrinsic alignment. Extensive experiments on diverse indoor and outdoor benchmarks demonstrate state-of-the-art accuracy in both pose estimation and scene reconstruction. Notably, M^3 reduces ATE RMSE by 64.3% compared to VGGT-SLAM 2.0 and outperforms ARTDECO by 2.11 dB in PSNR on the ScanNet++ dataset.

单目重建高斯泼溅多视图模型SLAM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。