arXiv:2503.07520cs.CVcs.IR2025-03中稿 · IEEE Transactions …被引 2

用少量标注数据实现无人机地理定位跨域迁移,解决无监督方法误标问题。

From Limited Labels to Open Domains:An Efficient Learning Method for Drone-view Geo-Localization

  • 构建双子网架构,学习跨视角结构与空间不变性特征
  • 仅需少量配对数据即在新域达到领先性能,少样本下超越现有无监督方法
  • 适合低标注成本、跨区域部署的无人机定位场景

传统监督式无人机地理定位方法严重依赖成对训练数据,在新领域部署时需重新采集配对数据并重训,计算开销大。现有无监督方法基于跨视图相似性生成伪标签,但地理相似性和空间连续性导致不同位置视觉特征相近,引发特征混淆,错误伪标签导致模型负优化。针对此问题,本文提出一种有限监督下的跨域不变知识迁移网络(CDIKTNet),包含跨域不变子网(CDIS)和跨域传输子网(CDTS)。CDIS利用少量配对数据学习跨视角结构与空间不变性,初始化未配对数据共享特征空间时即具备隐含的跨视图关联,缓解特征混淆。在此基础上,CDTS采用双路径对比学习进一步优化各子空间,同时保持共享空间一致性。大量实验表明,CDIKTNet在全监督下性能达当前最优,且在少样本与跨域初始化条件下均优于现有无监督方法。

原文摘要 · Abstract (English)

Traditional supervised drone-view geo-localization (DVGL) methods heavily depend on paired training data and encounter difficulties in learning cross-view correlations from unpaired data. Moreover, when deployed in a new domain, these methods require obtaining the new paired data and subsequent retraining for model adaptation, which significantly increases computational overhead. Existing unsupervised methods have enabled to generate pseudo-labels based on cross-view similarity to infer the pairing relationships. However, geographical similarity and spatial continuity often cause visually analogous features at different geographical locations. The feature confusion compromises the reliability of pseudo-label generation, where incorrect pseudo-labels drive negative optimization. Given these challenges inherent in both supervised and unsupervised DVGL methods, we propose a novel cross-domain invariant knowledge transfer network (CDIKTNet) with limited supervision, whose architecture consists of a cross-domain invariance sub-network (CDIS) and a cross-domain transfer sub-network (CDTS). This architecture facilitates a closed-loop framework for invariance feature learning and knowledge transfer. The CDIS is designed to learn cross-view structural and spatial invariance from a small amount of paired data that serves as prior knowledge. It endows the shared feature space of unpaired data with similar implicit cross-view correlations at initialization, which alleviates feature confusion. Based on this, the CDTS employs dual-path contrastive learning to further optimize each subspace while preserving consistency in a shared feature space. Extensive experiments demonstrate that CDIKTNet achieves state-of-the-art performance under full supervision compared with those supervised methods, and further surpasses existing unsupervised methods in both few-shot and cross-domain initialization.

地理定位无人机少样本跨域迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。