融合RGB与法向图,提升无人机跨视角定位精度
JRN-Geo: A Joint Perception Network based on RGB and Normal images for Cross-view Geo-localization
- 双分支网络同时处理颜色与几何结构信息
- 在University-1652和SUES-200上达到当前最优性能
- 适合需要高精度视觉定位的无人机应用
跨视角地理定位在无人机定位与导航中至关重要。然而,图像间视角差异大、外观变化显著,带来巨大挑战。现有方法主要依赖RGB图像的语义特征,常忽略空间结构信息对视角不变特征的贡献。为此,本文引入法向图中的几何结构信息,提出联合感知网络JRN-Geo,通过双分支特征提取框架,结合差异感知融合模块(DAFM)与联合约束交互聚合策略(JCIA),实现语义与结构信息的深度融合与联合约束表示。此外,提出3D地理增强技术生成潜在视角变化样本,增强网络学习视角不变特征的能力。在University-1652和SUES-200数据集上的大量实验验证了方法对复杂视角变化的鲁棒性,性能达当前最优。
原文摘要 · Abstract (English)
Cross-view geo-localization plays a critical role in Unmanned Aerial Vehicle (UAV) localization and navigation. However, significant challenges arise from the drastic viewpoint differences and appearance variations between images. Existing methods predominantly rely on semantic features from RGB images, often neglecting the importance of spatial structural information in capturing viewpoint-invariant features. To address this issue, we incorporate geometric structural information from normal images and introduce a Joint perception network to integrate RGB and Normal images (JRN-Geo). Our approach utilizes a dual-branch feature extraction framework, leveraging a Difference-Aware Fusion Module (DAFM) and Joint-Constrained Interaction Aggregation (JCIA) strategy to enable deep fusion and joint-constrained semantic and structural information representation. Furthermore, we propose a 3D geographic augmentation technique to generate potential viewpoint variation samples, enhancing the network's ability to learn viewpoint-invariant features. Extensive experiments on the University-1652 and SUES-200 datasets validate the robustness of our method against complex viewpoint ariations, achieving state-of-the-art performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。