arXiv:2511.20343cs.CV2025-11被引 12

用轻量体积表示实现高精度3D重建,无需微调即可适配多种视觉任务。

AMB3R: Accurate Feed-forward Metric-scale 3D Reconstruction with Backend

  • 采用紧凑稀疏体素作为后端,支持高效几何推理。
  • 在相机位姿、深度和度量尺度估计上达到领先水平。
  • 适合需要快速部署的3D重建与导航系统应用。

我们提出AMB3R,一种多视角前馈式模型,用于度量尺度下的密集3D重建,可应对多样化的3D视觉任务。核心思想是利用稀疏但紧凑的体素场景表示作为后端,实现空间紧凑的几何推理。尽管仅针对多视角重建训练,我们证明AMB3R可无缝扩展至非标定视觉里程计(在线)或大规模运动恢复结构,无需任务特定微调或测试时优化。相比先前基于点云的模型,本方法在相机位姿、深度及度量尺度估计、3D重建等方面均达当前最优性能,并在常见基准上超越基于优化的SLAM与SfM方法,即使在使用密集重建先验的情况下也表现更优。

原文摘要 · Abstract (English)

We present AMB3R, a multi-view feed-forward model for dense 3D reconstruction on a metric-scale that addresses diverse 3D vision tasks. The key idea is to leverage a sparse, yet compact, volumetric scene representation as our backend, enabling geometric reasoning with spatial compactness. Although trained solely for multi-view reconstruction, we demonstrate that AMB3R can be seamlessly extended to uncalibrated visual odometry (online) or large-scale structure from motion without the need for task-specific fine-tuning or test-time optimization. Compared to prior pointmap-based models, our approach achieves state-of-the-art performance in camera pose, depth, and metric-scale estimation, 3D reconstruction, and even surpasses optimization-based SLAM and SfM methods with dense reconstruction priors on common benchmarks.

3D重建视觉里程计体素表示前端模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。