arXiv:2604.17190cs.CV2026-04中稿 · CVPR被引 5

利用语言中的方向线索提升无人机导航准确性和效率

LookasideVLN: Direction-Aware Aerial Vision-and-Language Navigation

论文配图:LookasideVLN: Direction-Aware Aerial Vision-and-Language Navigation
图 1 · 摘自论文原文
  • 构建动态方向图,关联地标与方向信息
  • 单级前瞻下超越现有最佳方法,准确率显著提升
  • 适合研究无人机导航与多模态理解的学者

空中视觉语言导航(Aerial VLN)使无人机能依据自然语言指令在复杂城市环境中自主导航。尽管近期方法通过大规模记忆图和前瞻路径规划取得进展,但仍受限于对指令理解浅显及计算成本高,尤其依赖地标描述而忽略方向线索——人类导航中关键的空间上下文信息。本文提出LookasideVLN新范式,利用语言中的方向线索实现更精准的空间推理与更高计算效率。其包含三部分:(1) 自我中心方向图(ELG),动态编码与指令相关的地标及其方向关系;(2) 空间地标知识库(SLKB),从过往经验中轻量级检索记忆;(3) 方向前瞻多模态大模型导航代理(Lookaside MLLM),融合用户指令、视觉观测与ELG中的地標-方向信息进行路径规划。大量实验表明,即使仅使用单级前瞻,LookasideVLN仍显著优于当前最优的CityNavAgent,证明利用方向线索是高效且强大的Aerial VLN策略。

原文摘要 · Abstract (English)

Aerial Vision-and-Language Navigation (Aerial VLN) enables unmanned aerial vehicles (UAVs) to follow natural language instructions and navigate complex urban environments. While recent advances have achieved progress through large-scale memory graphs and lookahead path planning, they remain limited by shallow instruction understanding and high computational cost. In particular, existing methods rely primarily on landmark descriptions, overlooking directional cues "a key source of spatial context in human navigation". In this work, we propose LookasideVLN, a new paradigm that exploits directional cues in natural language to achieve both more accurate spatial reasoning and greater computational efficiency. LookasideVLN comprises three core components: (1) an Egocentric Lookaside Graph (ELG) that dynamically encodes instruction-relevant landmarks and their directional relationships, (2) a Spatial Landmark Knowledge Base (SLKB) that provides lightweight memory retrieval from prior navigation experiences, and (3) a Lookaside MLLM Navigation Agent that aligns multimodal information from user instructions, visual observations, and landmark-direction information from ELG for path planning. Extensive experiments show that LookasideVLN significantly outperforms the state-of-the-art CityNavAgent, even with a single-level lookahead, demonstrating that leveraging directional cues is a powerful yet efficient strategy for Aerial VLN.

无人机导航视觉语言方向感知多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。