用卫星图解决3D重建模型的尺度模糊问题。
Empowering Feed-Forward Reconstruction Models with Metric Scale via Satellite Images

- 利用卫星图像作为全局度量参考,通过双向跨视图交互校准
- 在KITTI、nuScenes等数据集上提升度量深度估计精度
- 适合需要真实尺度重建的自动驾驶与地图构建场景
前馈式3D重建模型虽在多种场景中表现出强泛化能力,但大多仅能恢复未知全局尺度的几何结构,限制了其在需环境度量理解的应用中的使用。现有度量重建方法通常依赖大规模度量标注或精确相机标定,而这些在多数实际场景中成本高或不可靠。本文提出一种卫星引导的框架,以解决前馈3D重建中的尺度模糊问题。核心思想是利用易获取的卫星影像作为全局度量参考。给定粗略相机位姿,方法检索局部卫星图像块,并通过双向跨视图交互将之与前馈重建主干融合。通过强制重建场景与卫星参考的一致性,模型推断出绝对尺度,优化场景几何,并在度量坐标系中估计相机位姿。在KITTI、nuScenes和Oxford RobotCar上的实验表明,该方法在度量深度估计、多视角点云重建和跨视图相机定位任务中均有持续提升,同时保持了跨数据集与地理区域的强泛化能力。
原文摘要 · Abstract (English)
Feed-forward 3D reconstruction models have recently shown strong generalization across diverse scenes, yet most of them recover geometry only up to an unknown global scale. This scale ambiguity limits their use in applications that require metric understanding of the environment. Existing metric reconstruction methods commonly rely on large-scale metric annotations or accurate camera calibration, both of which are costly or unreliable in many real-world settings. We propose a satellite-guided framework for resolving scale ambiguity in feed-forward 3D reconstruction. The key idea is to use readily available satellite imagery as a global metric reference. Given a coarse camera pose, our method retrieves a local satellite patch and integrates it with a feed-forward reconstruction backbone through bidirectional cross-view interaction. By enforcing consistency between the reconstructed scene and the satellite reference, the model infers absolute scale, refines scene geometry, and estimates camera pose in a metric coordinate frame. Experiments on KITTI, nuScenes, and Oxford RobotCar show consistent improvements in metric depth estimation, multi-view point-cloud reconstruction, and cross-view camera localization, while preserving strong generalization across datasets and geographic regions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。