arXiv:2504.20496cs.CV2025-04

提升野外视频的3D重建精度,解决动态物体与低纹理问题。

Large-scale visual SLAM for in-the-wild videos

  • 通过自监督恢复相机参数并用深度预测掩蔽动态区域
  • 结合单目深度与全局优化,实现长序列稳定重建
  • 适合机器人在复杂真实场景中快速部署使用

从非受控的野外视频中实现准确可靠的3D场景重建,可大幅简化机器人在新环境中的部署。然而,现有视觉仅SLAM方法在真实视频中表现不佳,常因快速旋转、纯平移、无纹理区域及动态物体导致失败。本文分析现有方法局限,提出一种鲁棒性更强的重建流水线:利用结构-从-运动自动恢复初始相机内参;采用预测模型掩蔽动态物体与弱约束区域;引入单目深度估计正则化束调整,缓解低视差情况下的误差;结合位置识别与回环检测,通过全局束调整减少长期漂移并优化内参与位姿。我们在多个在线视频上实现了大规模连续3D模型构建。相比基线方法常产生局部不一致结果(如分段或畸变地图),本系统在无真值位姿的情况下,通过重渲染NeRF模型的视觉一致性、执行时间和重建质量评估,展现出更长序列、更一致的重建效果,建立了当前野外视频视觉重建的新基准。

原文摘要 · Abstract (English)

Accurate and robust 3D scene reconstruction from casual, in-the-wild videos can significantly simplify robot deployment to new environments. However, reliable camera pose estimation and scene reconstruction from such unconstrained videos remains an open challenge. Existing visual-only SLAM methods perform well on benchmark datasets but struggle with real-world footage which often exhibits uncontrolled motion including rapid rotations and pure forward movements, textureless regions, and dynamic objects. We analyze the limitations of current methods and introduce a robust pipeline designed to improve 3D reconstruction from casual videos. We build upon recent deep visual odometry methods but increase robustness in several ways. Camera intrinsics are automatically recovered from the first few frames using structure-from-motion. Dynamic objects and less-constrained areas are masked with a predictive model. Additionally, we leverage monocular depth estimates to regularize bundle adjustment, mitigating errors in low-parallax situations. Finally, we integrate place recognition and loop closure to reduce long-term drift and refine both intrinsics and pose estimates through global bundle adjustment. We demonstrate large-scale contiguous 3D models from several online videos in various environments. In contrast, baseline methods typically produce locally inconsistent results at several points, producing separate segments or distorted maps. In lieu of ground-truth pose data, we evaluate map consistency, execution time and visual accuracy of re-rendered NeRF models. Our proposed system establishes a new baseline for visual reconstruction from casual uncontrolled videos found online, demonstrating more consistent reconstructions over longer sequences of in-the-wild videos than previously achieved.

视觉SLAM3D重建单目深度动态遮挡

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。