arXiv:2607.04352cs.CV2026-07被引 3

用扩散模型提升无人机最后10米精准导航能力

Last-Meter Precision Navigation for UAVs: A Diffusion-Refined Aerial Visual Servoing Approach

论文配图:Last-Meter Precision Navigation for UAVs: A Diffusion-Refined Aerial Visual Servoing Approach
图 1 · 摘自论文原文
  • 分两阶段:先用三角函数建模旋转,再用预训练世界模型模拟视觉未来
  • 在72个场景中实现零样本迁移,最后10米定位误差显著降低
  • 适合做无人机精准降落、复杂环境自主导航的研究者

本文研究无人机在最后一米范围内的精准导航问题,即仅用单目视觉实现最终10米内的自主抵达。该任务因尺度模糊性、旋转不连续性和精细空间推理需求而极具挑战,现有方法在视角剧烈变化下表现不佳,且泛化能力有限。为此,我们提出DreamNav——一种从粗到精的扩散优化空中视觉伺服框架。第一阶段采用三角函数参数化回归策略,联合建模正弦与余弦分量,有效缓解角度周期性带来的优化不稳定问题。基于此粗估计结果,第二阶段利用预训练世界模型模拟候选动作的未来视觉观测,通过视觉想象过程选择使视觉差异最小化的轨迹。为支持严格评估,我们构建了PairUAV基准数据集,包含来自University-1652数据集的480万对图像,覆盖72个场景。大量实验表明,DreamNav在精度和泛化性能上均优于主流视觉伺服与基础模型基线,具备零样本迁移到未见场景的能力。

原文摘要 · Abstract (English)

In this work, we study the last-meter precision navigation for UAVs, e.g., autonomously reaching a target within the final 10 meters using monocular vision. This task is challenging due to scale ambiguity, rotation discontinuities, and the need for fine-grained spatial reasoning. Existing methods often fail under large viewpoint changes or lack generalization to unseen environments. To this end, we propose DreamNav, a coarse-to-fine diffusion-refined aerial visual servoing framework. In the first coarse-estimation stage, a robust regression policy employs a trigonometric parameterization to predict rotation by jointly modeling sine and cosine components, effectively mitigating optimization instabilities caused by angular periodicity. Given this coarse estimate, the second diffusion-refined stage utilizes a pre-trained world model to simulate future visual observations for candidate actions, selecting the trajectory that minimizes visual discrepancy with the target through a process of visual imagination. To support rigorous evaluation, we contribute PairUAV, a large-scale benchmark comprising 4.8 million image pairs across 72 scenes, curated from the University-1652 dataset. Extensive experiments show DreamNav outperforms strong visual servoing and foundation model baselines in accuracy and generalization, with zero-shot transfer to unseen scenes.

无人机导航视觉伺服扩散模型精准定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。