分离内容与视角因素,提升无人机与卫星图像定位鲁棒性
Robust Drone-View Geo-Localization via Content-Viewpoint Disentanglement
- 将跨视角图像特征建模为内容与视角共同决定的复合流形
- 通过互信息最小化和跨视图重建约束实现显式解耦
- 可插拔集成,适用于多种场景且降低推理延迟
无人机视角地理定位(DVGL)旨在匹配从无人机和卫星视角拍摄的同一地理位置图像。尽管近期取得进展,但由于视角变化带来的显著外观差异和空间畸变,该任务仍具挑战性。现有方法通常假设通过对比学习可在共享特征空间中直接对齐无人机与卫星图像,但这一假设忽略了视角差异引发的本质冲突,导致提取的特征包含不一致信息,阻碍精确定位。本文从流形学习角度出发,将跨视角图像的特征空间建模为由内容与视角共同支配的复合流形。基于此,提出新框架CVD,显式解耦内容与视角因子。为促进有效解耦,引入两项约束:(i) 视图内独立性约束,通过最小化两因子间的互信息以实现统计独立;(ii) 视图间重建约束,利用配对图像中的内容与视角信息交叉重构各视图,确保因子语义专属性。作为即插即用模块,CVD可无缝融入现有DVGL流程,降低推理延迟。在University-1652和SUES-200上的大量实验表明,CVD在多种场景、视角及高度下表现出强鲁棒性与泛化能力;在CVUSA和CVACT上的进一步评估也证实其持续提升效果。
原文摘要 · Abstract (English)
Drone-view geo-localization (DVGL) aims to match images of the same geographic location captured from drone and satellite perspectives. Despite recent advances, DVGL remains challenging due to significant appearance changes and spatial distortions caused by viewpoint variations. Existing methods typically assume that drone and satellite images can be directly aligned in a shared feature space via contrastive learning. Nonetheless, this assumption overlooks the inherent conflicts induced by viewpoint discrepancies, resulting in extracted features containing inconsistent information that hinders precise localization. In this study, we take a manifold learning perspective and model $\textit{the feature space of cross-view images as a composite manifold jointly governed by content and viewpoint}$. Building upon this insight, we propose $\textbf{CVD}$, a new DVGL framework that explicitly disentangles $\textit{content}$ and $\textit{viewpoint}$ factors. To promote effective disentanglement, we introduce two constraints: $\textit{(i)}$ an intra-view independence constraint that encourages statistical independence between the two factors by minimizing their mutual information; and $\textit{(ii)}$ an inter-view reconstruction constraint that reconstructs each view by cross-combining $\textit{content}$ and $\textit{viewpoint}$ from paired images, ensuring factor-specific semantics are preserved. As a plug-and-play module, CVD integrates seamlessly into existing DVGL pipelines and reduces inference latency. Extensive experiments on University-1652 and SUES-200 show that CVD exhibits strong robustness and generalization across various scenarios, viewpoints and altitudes, with further evaluations on CVUSA and CVACT confirming consistent improvements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。