无需标注数据,通过迭代渲染实现无人机图像跨视角地理定位
Unsupervised Multi-view UAV Image Geo-localization via Iterative Rendering
- 从无人机图像构建三维场景,生成类卫星视图以减少视角差异
- 在University-1652和SUES-200上达到与有监督方法相当的定位精度
- 适合无标注数据、需跨区域泛化的无人机地理定位任务
无人机跨视角地理定位因倾斜航拍图与俯视卫星图之间的视角差异而面临挑战。现有方法依赖标注数据提取视角不变特征,但训练成本高且易过拟合局部区域特征,泛化能力弱。为此,我们提出一种无监督方案:从无人机观测中提升场景表示至3D空间,生成类卫星视图,增强对视角失真的鲁棒性。通过生成接近卫星视角的正交图像,降低特征表示中的视角差异,缓解区域特定图像匹配的捷径问题。为进一步对齐渲染图像与真实卫星视角,设计了迭代相机位姿更新机制,逐步调节渲染查询图与潜在卫星目标的对齐,消除相对于参考图像的空间偏移。该迭代优化策略通过多轮视图一致性融合,强化跨视角特征不变性。因此,该无监督范式天然避免区域过拟合问题,无需特征微调或数据驱动训练即可实现通用的无人机图像地理定位。在University-1652和SUES-200数据集上的实验表明,本方法显著提升定位准确率,并在多样区域间保持鲁棒性。值得注意的是,不进行模型微调或成对训练,仍可达到近期有监督方法的竞争力。
原文摘要 · Abstract (English)
Unmanned Aerial Vehicle (UAV) Cross-View Geo-Localization (CVGL) presents significant challenges due to the view discrepancy between oblique UAV images and overhead satellite images. Existing methods heavily rely on the supervision of labeled datasets to extract viewpoint-invariant features for cross-view retrieval. However, these methods have expensive training costs and tend to overfit the region-specific cues, showing limited generalizability to new regions. To overcome this issue, we propose an unsupervised solution that lifts the scene representation to 3d space from UAV observations for satellite image generation, providing robust representation against view distortion. By generating orthogonal images that closely resemble satellite views, our method reduces view discrepancies in feature representation and mitigates shortcuts in region-specific image pairing. To further align the rendered image's perspective with the real one, we design an iterative camera pose updating mechanism that progressively modulates the rendered query image with potential satellite targets, eliminating spatial offsets relative to the reference images. Additionally, this iterative refinement strategy enhances cross-view feature invariance through view-consistent fusion across iterations. As such, our unsupervised paradigm naturally avoids the problem of region-specific overfitting, enabling generic CVGL for UAV images without feature fine-tuning or data-driven training. Experiments on the University-1652 and SUES-200 datasets demonstrate that our approach significantly improves geo-localization accuracy while maintaining robustness across diverse regions. Notably, without model fine-tuning or paired training, our method achieves competitive performance with recent supervised methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。