arXiv:2504.09587cs.RO2025-04被引 19

让无人机听懂地图指令,精准找到城市目标

GeoNav: Empowering MLLMs with dual-scale geospatial reasoning for language-goal aerial navigation

  • 分三阶段模拟人类粗到细的空间推理,动态构建全局地图与局部场景图
  • 在CityNav上成功率达84.3%,比当前最优提升18.4%
  • 适合研究无人机导航、多模态大模型与地理空间理解的学者

语言目标空中导航要求无人机在复杂户外环境(如城市街区)中根据文本指令定位目标。现有室内方法难以拓展至城市场景,因物体模糊、视野受限及空间推理困难。本文提出GeoNav,一种具备地理空间感知能力的多模态智能体,支持长距离空中导航。其运行分为三个阶段:地标导航、目标搜索与精确定位,模仿人类由粗到细的空间推理模式。为支撑该推理,动态构建双尺度空间表征:一是融合先验地理知识与视觉线索的全局概略认知地图,以俯视形式显式标注,实现基于地图的快速导航;二是表达地标与物体间层级关系的局部精细场景图,用于精准定位目标。在此结构化记忆基础上,引入空间思维链机制,使多模态大模型能高效、可解释地跨阶段决策。在CityNav基准测试中,GeoNav成功率达84.3%,较当前最优提升18.4%,显著降低导航误差。消融实验表明各模块均关键,结构化空间感知是高级无人机导航的核心。

原文摘要 · Abstract (English)

Language-goal aerial navigation requires UAVs to localize targets in the complex outdoors, such as urban blocks based on textual instructions. The indoor methods are often hard to scale to urban scenes due to ambiguous objects, limited visual field, and spatial reasoning. In this work, we propose GeoNav, a multi-modal agent for long-range aerial navigation with geospatial awareness. GeoNav operates in three phases-landmark navigation, target search, and precise localization-mimicking human coarse-to-fine spatial reasoning patterns. To support such reasoning, it dynamically builds dual-scale spatial representations. The first is a global but schematic cognitive map, which fuses prior geographic knowledge and embodied visual cues into a top-down and explicit annotated form. It enables fast navigation to the landmark region via intuitive map-based reasoning. The second is a local but delicate scene graph representing hierarchical spatial relationships between landmarks and objects, utilized for accurate target localization. On top of the structured memory, GeoNav employs a spatial chain-of-thought mechanism to enable MLLMs with efficient and interpretable action-making across stages. On the CityNav benchmark, GeoNav surpasses the current SOTA up to 18.4% in success rate and significantly eliminates navigation error. The ablation studies highlight the importance of each module, positioning structured spatial perception as the key to advanced UAV navigation. Published in Pattern Recognition, 2026.

无人机导航多模态模型空间推理地理感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。