arXiv:2604.07705cs.RO2026-04被引 2

综述无人机视觉语言导航新进展,聚焦大模型融合与未来挑战

Vision-Language Navigation for Aerial Robots: Towards the Era of Large Language Models

  • 按架构分类五类导航方法,分析设计原理与性能差异
  • 指出数据集规模小、环境单一、评估指标不足等关键缺陷
  • 提出7个未解难题,如长指令理解、多机协同与真实部署

空中视觉语言导航(Aerial VLN)旨在使无人飞行器(UAV)理解自然语言指令,在复杂三维环境中通过视觉感知实现自主导航。本文对这一领域进行批判性综述,重点分析大语言模型(LLMs)与视觉语言模型(VLMs)的最新融合趋势。首先明确Aerial VLN问题,定义单指令与对话式两种交互范式。将现有方法归纳为五类:序列到序列与注意力机制方法、端到端 LLM/VLM 方法、层次化方法、多智能体方法及对话式导航方法。对每类方法系统分析设计动机、技术权衡与实测表现。批判性评估当前评价体系,包括数据集、仿真平台与评测指标,揭示其在规模、环境多样性、现实锚定与度量覆盖方面的不足。在共享基准上进行跨方法比较,分析离散与连续动作、端到端与分层结构、仿真到现实差距等核心权衡。最终提炼出七个具体开放问题:长时序指令理解、视角鲁棒性、可扩展空间表征、连续6-自由度动作执行、机载部署、基准标准化及多无人机集群导航,并基于证据提出具体研究方向。

原文摘要 · Abstract (English)

Aerial vision-and-language navigation (Aerial VLN) aims to enable unmanned aerial vehicles (UAVs) to interpret natural language instructions and autonomously navigate complex three-dimensional environments by grounding language in visual perception. This survey provides a critical and analytical review of the Aerial VLN field, with particular attention to the recent integration of large language models (LLMs) and vision-language models (VLMs). We first formally introduce the Aerial VLN problem and define two interaction paradigms: single-instruction and dialog-based, as foundational axes. We then organize the body of Aerial VLN methods into a taxonomy of five architectural categories: sequence-to-sequence and attention-based methods, end-to-end LLM/VLM methods, hierarchical methods, multi-agent methods, and dialog-based navigation methods. For each category, we systematically analyze design rationales, technical trade-offs, and reported performance. We critically assess the evaluation infrastructure for Aerial VLN, including datasets, simulation platforms, and metrics, and identify their gaps in scale, environmental diversity, real-world grounding, and metric coverage. We consolidate cross-method comparisons on shared benchmarks and analyze key architectural trade-offs, including discrete versus continuous actions, end-to-end versus hierarchical designs, and the simulation-to-reality gap. Finally, we synthesize seven concrete open problems: long-horizon instruction grounding, viewpoint robustness, scalable spatial representation, continuous 6-DoF action execution, onboard deployment, benchmark standardization, and multi-UAV swarm navigation, with specific research directions grounded in the evidence presented throughout the survey.

无人机导航大模型视觉语言综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。