构建跨空地视觉定位数据集,解决城市环境定位难题
Aerial-ground Cross-modal Localization: Dataset, Ground-truth, and Benchmark
- 融合车载影像与机载激光点云,构建跨平台数据集
- 覆盖武汉、香港、旧金山三座城市,规模超百万级图像
- 为城市级视觉定位提供真实标注基准,适合自动驾驶研究
在密集城市环境中实现精确视觉定位是摄影测量学、地理空间信息科学和机器人领域的重要任务。尽管图像作为低成本且广泛可用的感知模态,其在视觉里程计中的表现常受限于无纹理表面、剧烈视角变化和长期漂移。随着机载激光扫描(ALS)数据的公开获取,利用ALS作为先验地图可为大规模高精度视觉定位开辟新路径。然而,由于三大局限:(1)缺乏多平台数据集,(2)缺少适用于大规模城市环境的可靠真值生成方法,(3)现有图像到点云(I2P)算法在空-地跨平台设置下的验证不足,基于ALS的定位潜力尚未充分挖掘。为此,我们引入一个大规模数据集,整合了武汉、香港和旧金山地区移动测绘系统采集的地面影像与对应的机载激光点云。
原文摘要 · Abstract (English)
Accurate visual localization in dense urban environments poses a fundamental task in photogrammetry, geospatial information science, and robotics. While imagery is a low-cost and widely accessible sensing modality, its effectiveness on visual odometry is often limited by textureless surfaces, severe viewpoint changes, and long-term drift. The growing public availability of airborne laser scanning (ALS) data opens new avenues for scalable and precise visual localization by leveraging ALS as a prior map. However, the potential of ALS-based localization remains underexplored due to three key limitations: (1) the lack of platform-diverse datasets, (2) the absence of reliable ground-truth generation methods applicable to large-scale urban environments, and (3) limited validation of existing Image-to-Point Cloud (I2P) algorithms under aerial-ground cross-platform settings. To overcome these challenges, we introduce a new large-scale dataset that integrates ground-level imagery from mobile mapping systems with ALS point clouds collected in Wuhan, Hong Kong, and San Francisco.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。