用无人机图像辅助卫星与地面图像,实现更精准的三维定位与重建。
Seeing Across Skies and Streets: Feedforward 3D Reconstruction from Satellite, Drone, and Ground Images

- 引入无人机图像作为中间视角,突破传统二维定位限制。
- 单次前向传播即可恢复6自由度相机位姿与跨视图点云。
- 适用于真实地形中的复杂场景,适合自动驾驶与地理信息领域。
跨视图定位经典问题为:给定一张地面图像,它在卫星影像中的位置在哪里?现有方法通常仅能估计3自由度(即(x,y)位置和偏航角),因垂直俯拍的卫星图像无法提供滚转、俯仰或高度的直接线索,导致依赖平面运动与零倾斜假设。这些假设在有坡度、斜坡或倾斜镜头的真实地形中失效。为此,本文引入单一无人机图像作为中间视角:它揭示了从正上方无法获取的三维结构,提供了卫星图像缺失的滚转、俯仰和高度线索,且只需与地面相机存在空间重叠——无需已知相对位姿。基于此洞察,我们提出**Cross3R**,一种灵活的前馈模型,可接收卫星影像块、无人机图像、地面图像或其组合,在一次前向传播中恢复跨视图三维点云、所有输入相机的6-DoF位姿,以及每张图像在卫星图块上的(x,y)位置和偏航角。为训练与评估,我们构建了**CrossGeo**数据集,包含27.8万张图像,覆盖85个场景,遍布除南极洲外的所有大陆。在CrossGeo上,Cross3R在点云重建、6-DoF相机位姿估计和跨视图定位任务中持续优于现有前馈3D基线模型;在KITTI上,尽管未使用任何KITTI训练数据,其性能仍优于专门在KITTI上训练的跨视图方法,多数指标领先。
原文摘要 · Abstract (English)
Cross-view localization classically asks: where does this ground image lie on the satellite tile? Existing methods are typically limited to 3-DoF estimates -- an $(x,y)$ position and a yaw angle -- because nadir satellite imagery provides no direct cues for roll, pitch, or altitude, forcing a reliance on planar-motion and zero-tilt assumptions. These assumptions break on real terrain with slopes, ramps, and tilted camera mounts. To overcome this, we introduce a single UAV image as an intermediate viewpoint: it reveals the 3D structure invisible from nadir, supplies the cues for roll, pitch, and altitude that the satellite alone cannot provide, and needs only spatial overlap with the ground camera -- no known relative pose is required. Building on this insight, we propose **Cross3R**, a flexible feed-forward model that ingests a satellite tile together with a UAV image, a ground image, or both, and, in a single forward pass, recovers a cross-view 3D point cloud, the 6-DoF poses of every input camera, and the on-tile $(x,y)$ position and yaw of each perspective camera. For training and evaluation, we also construct **CrossGeo**, a 278K-image tri-view dataset spanning 85 scenes across every continent except Antarctica. On CrossGeo, Cross3R consistently outperforms feed-forward 3D baselines in point-cloud reconstruction, 6-DoF camera-pose estimation, and cross-view localization. On KITTI, it outperforms dedicated cross-view methods trained on KITTI on most metrics, despite having no KITTI training itself.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。