arXiv:2506.10756cs.ROcs.AI2025-06被引 23

让无人机通过语言指令自主飞行,支持抽象指令和复杂环境导航

Grounded Vision-Language Navigation for UAVs with Open-Vocabulary Goal Understanding

  • 用大模型重构指令,视觉语言模型匹配目标图像生成路径
  • 无需定位或测距传感器,仅靠单目摄像头实现连续速度控制
  • 在真实室内外环境验证,对抽象语言指令仍具强泛化能力

视觉-语言导航(VLN)是自主机器人领域长期挑战,旨在使智能体在复杂环境中根据人类指令完成导航。当前主要瓶颈在于对分布外环境的泛化能力不足,以及对固定离散动作空间的依赖。为此,我们提出专为无人机设计的Vision-Language Fly(VLFly)框架,实现语言引导的飞行任务。VLFly无需定位或主动测距传感器,仅通过机载单目相机捕捉的自我中心观测,输出连续速度指令。该框架包含三个模块:基于大语言模型(LLM)的指令编码器,将高层语言指令转化为结构化提示;由视觉语言模型(VLM)驱动的目标检索器,通过视觉-语言相似性匹配提示与目标图像;以及生成可执行轨迹的航点规划器,实现实时无人机控制。VLFly在多种仿真环境中无须额外微调即持续优于所有基线。此外,在真实室内外环境下,面对直接与间接指令,均展现出鲁棒的开放词汇目标理解与泛化导航能力,即使面对抽象语言输入也表现稳定。

原文摘要 · Abstract (English)

Vision-and-language navigation (VLN) is a long-standing challenge in autonomous robotics, aiming to empower agents with the ability to follow human instructions while navigating complex environments. Two key bottlenecks remain in this field: generalization to out-of-distribution environments and reliance on fixed discrete action spaces. To address these challenges, we propose Vision-Language Fly (VLFly), a framework tailored for Unmanned Aerial Vehicles (UAVs) to execute language-guided flight. Without the requirement for localization or active ranging sensors, VLFly outputs continuous velocity commands purely from egocentric observations captured by an onboard monocular camera. The VLFly integrates three modules: an instruction encoder based on a large language model (LLM) that reformulates high-level language into structured prompts, a goal retriever powered by a vision-language model (VLM) that matches these prompts to goal images via vision-language similarity, and a waypoint planner that generates executable trajectories for real-time UAV control. VLFly is evaluated across diverse simulation environments without additional fine-tuning and consistently outperforms all baselines. Moreover, real-world VLN tasks in indoor and outdoor environments under direct and indirect instructions demonstrate that VLFly achieves robust open-vocabulary goal understanding and generalized navigation capabilities, even in the presence of abstract language input.

无人机导航视觉语言开放词汇

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。