arXiv:2607.17171cs.RO2026-07

用视觉惯性数据锚定单目大模型的尺度,实现高精度三维重建。

VIDAR: Visual-Inertial Dense Alignment and Reconstruction via a Geometric Foundation Model

论文配图:VIDAR: Visual-Inertial Dense Alignment and Reconstruction via a Geometric Foundation Model
图 1 · 摘自论文原文
  • 融合视觉惯性里程计与深度一切3,以姿态约束对齐时序几何预测。
  • 在EuRoC数据集上,姿态注入使尺度误差降至1%,平均精度达0.463。
  • 无需真值姿态即可实现0.676的精度,适合移动设备实时重建应用。

单目基础模型虽能提供稠密几何信息,但通常缺乏稳定度量尺度。本文提出VIDAR,一种结合SVO+IMU里程计与深度一切3(DA3)的视觉-惯性稠密重建框架。该框架以视觉-惯性前端作为度量锚点,提供相机位姿、尺度及一致的世界坐标系,用于对齐不同时刻的基础模型预测。基础模型则贡献细节丰富的局部几何,并融合进全局重建中。研究对比了基于姿态条件的DA3与解耦对齐策略,在EuRoC数据集上,姿态注入将尺度误差降低至约1%,达到0.463的平均[email protected];解耦混合策略在无真值姿态条件下进一步提升至0.676。EuRoC与TUM RGB-D实验表明,VIDAR是实现度量稠密单目重建的实用路径。

原文摘要 · Abstract (English)

Monocular foundation models provide dense geometry but usually lack a stable metric scale. This paper presents VIDAR, a visual-inertial dense reconstruction framework that couples SVO+IMU odometry with Depth Anything 3. VIDAR uses the visual-inertial front end as a metric anchor: it provides camera poses, scale, and a consistent world frame for aligning dense foundation-model predictions across time. The foundation model then contributes detailed local geometry that is fused into a global reconstruction. We study both pose-conditioned DA3 and a decoupled alignment strategy. On EuRoC, pose injection reduces scale error to about 1\% and reaches 0.463 mean [email protected]; the decoupled hybrid improves this to 0.676 without ground-truth poses. Results on EuRoC and TUM RGB-D show that VIDAR is a practical route to metric dense monocular reconstruction.

三维重建视觉惯性基础模型单目估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。