arXiv:2602.04517cs.CVcs.RO2026-02被引 3

用滑动窗口提升单目3D重建的可扩展性,无需重训练

S-MUSt3R: Sliding Multi-view 3D Reconstruction

  • 通过分段处理+对齐+轻量闭环优化突破大序列重建内存瓶颈
  • 在TUM、7-Scenes等数据集上实现与复杂方法相当的重建精度
  • 适合需要实时、高精度度量空间重建的机器人导航场景

近期3D视觉范式转向基础模型,其在未标定图像下的3D感知能力显著。然而,将此类模型扩展至大规模RGB流3D重建仍受内存限制。本文提出S-MUSt3R,一种简单高效的单目3D重建流水线,通过序列分段、分段对齐及轻量闭环优化,突破基础模型的可扩展性瓶颈。无需模型重训练,即可利用MUSt3R模型强大的3D重建能力,在TUM、7-Scenes及自有机器人导航数据集上成功处理长视频序列,生成准确且一致的3D重建结果。实验表明,该方法在轨迹和重建性能上媲美传统复杂架构,且能直接输出度量空间预测,具备在真实场景中规模化应用的潜力。

原文摘要 · Abstract (English)

The recent paradigm shift in 3D vision led to the rise of foundation models with remarkable capabilities in 3D perception from uncalibrated images. However, extending these models to large-scale RGB stream 3D reconstruction remains challenging due to memory limitations. This work proposes S-MUSt3R, a simple and efficient pipeline that extends the limits of foundation models for monocular 3D reconstruction. Our approach addresses the scalability bottleneck of foundation models through a simple strategy of sequence segmentation followed by segment alignment and lightweight loop closure optimization. Without model retraining, we benefit from remarkable 3D reconstruction capacities of MUSt3R model and achieve trajectory and reconstruction performance comparable to traditional methods with more complex architecture. We evaluate S-MUSt3R on TUM, 7-Scenes and proprietary robot navigation datasets and show that S-MUSt3R runs successfully on long RGB sequences and produces accurate and consistent 3D reconstruction. Our results highlight the potential of leveraging the MUSt3R model for scalable monocular 3D scene in real-world settings, with an important advantage of making predictions directly in the metric space.

3D重建单目视觉机器人导航基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。