arXiv:2503.11091cs.CV2025-03被引 13

让无人机按指令飞行,通过网格视角选择实现上下左右协同导航

Aerial Vision-and-Language Navigation with Grid-based View Selection and Map Construction

  • 用网格化视角选择替代传统动作预测,显式建模垂直与水平动作耦合
  • 构建鸟瞰图地图融合历史视觉信息,有效缓解障碍物干扰
  • 适合需要精准三维空间导航的无人机任务,尤其复杂空域场景

空中视觉-语言导航旨在使无人飞行器根据人类指令在三维空中环境自主导航。相较于地面导航,空中导航需同时决策水平与垂直方向的动作,路径更长、三维场景更复杂,且以往方法忽视了垂直与水平动作间的相互影响。本文提出一种基于网格的视角选择框架,将空中导航动作预测转化为网格化视角选择任务,并以兼顾水平与垂直动作的方式进行高度调整,实现有效垂直控制。进一步引入基于网格的鸟瞰图地图,融合导航历史中的视觉信息,提供上下文场景知识并减轻障碍物影响。最后采用跨模态变换器,显式对齐长期导航历史与指令。大量实验验证了所提方法的优越性。

原文摘要 · Abstract (English)

Aerial Vision-and-Language Navigation (Aerial VLN) aims to obtain an unmanned aerial vehicle agent to navigate aerial 3D environments following human instruction. Compared to ground-based VLN, aerial VLN requires the agent to decide the next action in both horizontal and vertical directions based on the first-person view observations. Previous methods struggle to perform well due to the longer navigation path, more complicated 3D scenes, and the neglect of the interplay between vertical and horizontal actions. In this paper, we propose a novel grid-based view selection framework that formulates aerial VLN action prediction as a grid-based view selection task, incorporating vertical action prediction in a manner that accounts for the coupling with horizontal actions, thereby enabling effective altitude adjustments. We further introduce a grid-based bird's eye view map for aerial space to fuse the visual information in the navigation history, provide contextual scene information, and mitigate the impact of obstacles. Finally, a cross-modal transformer is adopted to explicitly align the long navigation history with the instruction. We demonstrate the superiority of our method in extensive experiments.

无人机导航视觉语言三维空间网格建图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。