用小车作为度量参照,解决无人机与卫星图像尺度不一致问题。
Scale-Aware UAV-to-Satellite Cross-View Geo-Localization: A Semantic Geometric Approach
- 利用小车的先验尺寸和检测特征,恢复单目无人机图像的绝对尺度。
- 在未知尺度下,跨视角定位准确率提升显著,尤其在真实复杂场景中。
- 适合需要高精度定位或无源测高的无人机应用,如城市监控、测绘等。
无人机图像与卫星图像之间的跨视图地理定位(CVGL)在目标定位和无人机自定位中至关重要。然而,现有方法多依赖于无人机查询与卫星画廊间尺度一致的理想假设,忽视了实际场景中普遍存在的尺度模糊问题。这种差异导致视场错位和特征不匹配,严重降低定位鲁棒性。为此,我们提出一种几何框架,通过语义锚点从单目无人机图像中恢复绝对度量尺度。具体地,利用小车(SVs)——其具有相对稳定的先验尺寸分布且易于检测——作为度量参考。引入解耦立体投影模型,从这些语义目标估计绝对图像尺度。通过将车辆尺寸分解为径向与切向分量,补偿三维车辆在二维检测中的透视畸变,实现更精确的尺度估计。为进一步降低类内尺寸差异和检测噪声,采用基于四分位距(IQR)的鲁棒聚合双维度融合策略。所估计的全局尺度被用于尺度自适应的卫星图像裁剪,提升无人机到卫星的特征对齐效果。在增强版DenseUAV和UAV-VisLoc数据集上的实验表明,该方法在未知无人机图像尺度条件下显著提升了CVGL的鲁棒性。此外,该框架在下游任务如被动式无人机高度估计和3D模型尺度恢复方面也展现出强潜力。
原文摘要 · Abstract (English)
Cross-View Geo-Localization (CVGL) between UAV imagery and satellite images plays a crucial role in target localization and UAV self-positioning. However, most existing methods rely on the idealized assumption of scale consistency between UAV queries and satellite galleries, overlooking the severe scale ambiguity commonly encountered in real-world scenarios. This discrepancy leads to field-of-view misalignment and feature mismatch, significantly degrading CVGL robustness. To address this issue, we propose a geometric framework that recovers the absolute metric scale from monocular UAV images using semantic anchors. Specifically, small vehicles (SVs), characterized by relatively stable prior size distributions and high detectability, are exploited as metric references. A Decoupled Stereoscopic Projection Model is introduced to estimate the absolute image scale from these semantic targets. By decomposing vehicle dimensions into radial and tangential components, the model compensates for perspective distortions in 2D detections of 3D vehicles, enabling more accurate scale estimation. To further reduce intra-class size variation and detection noise, a dual-dimension fusion strategy with Interquartile Range (IQR)-based robust aggregation is employed. The estimated global scale is then used as a physical constraint for scale-adaptive satellite image cropping, improving UAV-to-satellite feature alignment. Experiments on augmented DenseUAV and UAV-VisLoc datasets demonstrate that the proposed method significantly improves CVGL robustness under unknown UAV image scales. Additionally, the framework shows strong potential for downstream applications such as passive UAV altitude estimation and 3D model scale recovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。