arXiv:2511.06422cs.CV2025-11被引 3

用扩散模型实现无人机与卫星图像跨视角定位,无需文本提示

DiffusionUavLoc: Visually Prompted Diffusion for Cross-View UAV Localization

  • 用无训练几何渲染生成卫星视角伪图像作结构提示
  • 在University-1652上卫星到无人机定位准确率超基线方法
  • 不依赖文本、标注和复杂网络,适合实际部署场景

随着低空经济快速发展,无人机已成为智能巡检系统中测量与追踪的关键平台。但在无GNSS环境下,依赖卫星信号的定位方案易失效。基于跨视角图像检索的定位是可行替代方案,但倾斜无人机图像与正射卫星影像间存在显著几何与外观差异。传统方法常依赖复杂网络架构、文本提示或大量标注,限制了泛化能力。为此,我们提出DiffusionUavLoc,一种图像提示、无文本、以扩散模型为核心的跨视角定位框架,采用VAE实现统一表征。首先通过无训练几何渲染,从无人机图像合成伪卫星图像作为结构提示;随后设计无文本条件扩散模型,融合多模态结构线索,学习对视角变化鲁棒的特征表示。推理时,在固定时间步t计算描述符,并使用余弦相似度匹配。在University-1652和SUES-200数据集上,该方法在跨视角定位任务中表现优异,尤其在University-1652的卫星到无人机场景中性能领先。代码与数据将公开于https://github.com/liutao23/DiffusionUavLoc.git。

原文摘要 · Abstract (English)

With the rapid growth of the low-altitude economy, unmanned aerial vehicles (UAVs) have become key platforms for measurement and tracking in intelligent patrol systems. However, in GNSS-denied environments, localization schemes that rely solely on satellite signals are prone to failure. Cross-view image retrieval-based localization is a promising alternative, yet substantial geometric and appearance domain gaps exist between oblique UAV views and nadir satellite orthophotos. Moreover, conventional approaches often depend on complex network architectures, text prompts, or large amounts of annotation, which hinders generalization. To address these issues, we propose DiffusionUavLoc, a cross-view localization framework that is image-prompted, text-free, diffusion-centric, and employs a VAE for unified representation. We first use training-free geometric rendering to synthesize pseudo-satellite images from UAV imagery as structural prompts. We then design a text-free conditional diffusion model that fuses multimodal structural cues to learn features robust to viewpoint changes. At inference, descriptors are computed at a fixed time step t and compared using cosine similarity. On University-1652 and SUES-200, the method performs competitively for cross-view localization, especially for satellite-to-drone in University-1652.Our data and code will be published at the following URL: https://github.com/liutao23/DiffusionUavLoc.git.

跨视角定位扩散模型无人机导航

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。