arXiv:2609.08442cs.CVcs.AI2026-09

用空间锚点融合局部与全局信息,提升无人机零样本导航能力

AirAnchor: Bridging Local and Global Spatial Information for Zero-Shot Aerial Vision-and-Language Navigation

论文配图:AirAnchor: Bridging Local and Global Spatial Information for Zero-Shot Aerial Vision-and-Language Navigation
图 1 · 摘自论文原文
  • 通过视觉观测动态构建局部空间锚点,实现即时定位
  • 持续维护物体知识库作为全局记忆,支持长程路径规划
  • 适合需要跨尺度空间理解的无人机导航研究者

航空视觉-语言导航要求无人机根据自然语言指令在复杂城市环境中导航。准确导航依赖于局部和全局空间信息,分别支持即时动作定位和长期路径规划。然而,现有零样本方法通常仅在单一空间尺度上运行,要么依赖当前观测在线构建局部表示,要么依赖离线历史经验构建全局记忆。为此,我们提出AirAnchor,一种通过空间锚点连接局部与全局空间信息的新范式,并将其整合到统一导航框架中,实现全面的空间语义定位以支持决策。AirAnchor包含三个核心组件:(1) 查询驱动的空间锚点定位,从视觉观测中识别决策相关锚点并组织为局部空间表示;(2) 持久物体空间记忆,增量式维护物体知识库作为持久化全局记忆,并检索与地标相关的空间先验;(3) 空间感知导航代理,显式将局部与全局空间信息融入智能体框架进行决策。在AerialVLN上的大量实验表明,AirAnchor显著优于现有零样本基线,验证了该范式的有效性和高效性。

原文摘要 · Abstract (English)

Aerial Vision-and-Language Navigation requires drones to follow natural-language instructions and navigate through complex urban environments. Accurate navigation relies on both local and global spatial information, which support immediate action grounding and long-horizon path planning, respectively. However, existing zero-shot methods typically operate at a single spatial scale, relying either on local representations constructed online from current observations or on global memories built offline from historical experience. To address this limitation, we propose AirAnchor, a new paradigm that bridges local and global spatial information through spatial anchors and integrates both into a shared navigation framework, enabling comprehensive spatial grounding for decision-making. AirAnchor consists of three core components: (1) Query-Driven Spatial Anchor Grounding, which identifies decision-relevant anchors from visual observations and organizes them into local spatial representations; (2) Persistent Object Spatial Memory, which incrementally maintains an object knowledge base as persistent global spatial memory and retrieves landmark-related spatial priors; and (3) a Spatially-Informed Navigation Agent, which explicitly integrates both local and global spatial information into an agentic framework for decision-making. Extensive experiments on AerialVLN demonstrate that AirAnchor substantially outperforms existing zero-shot baselines, validating the effectiveness and efficiency of the proposed paradigm.

无人机导航多模态空间记忆零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。