综述无人机视觉语言导航新进展,聚焦大模型融合与未来挑战
Vision-Language Navigation for Aerial Robots: Towards the Era of Large Language Models
- 按架构分类五类导航方法,分析设计原理与性能差异
- 指出数据集规模小、环境单一、评估指标不足等关键缺陷
- 提出7个未解难题,如长指令理解、多机协同与真实部署
空中视觉语言导航(Aerial VLN)旨在使无人飞行器(UAV)理解自然语言指令,在复杂三维环境中通过视觉感知实现自主导航。本文对这一领域进行批判性综述,重点分析大语言模型(LLMs)与视觉语言模型(VLMs)的最新融合趋势。首先明确Aerial VLN问题,定义单指令与对话式两种交互范式。将现有方法归纳为五类:序列到序列与注意力机制方法、端到端 LLM/VLM 方法、层次化方法、多智能体方法及对话式导航方法。对每类方法系统分析设计动机、技术权衡与实测表现。批判性评估当前评价体系,包括数据集、仿真平台与评测指标,揭示其在规模、环境多样性、现实锚定与度量覆盖方面的不足。在共享基准上进行跨方法比较,分析离散与连续动作、端到端与分层结构、仿真到现实差距等核心权衡。最终提炼出七个具体开放问题:长时序指令理解、视角鲁棒性、可扩展空间表征、连续6-自由度动作执行、机载部署、基准标准化及多无人机集群导航,并基于证据提出具体研究方向。
原文摘要 · Abstract (English)
Aerial vision-and-language navigation (Aerial VLN) aims to enable unmanned aerial vehicles (UAVs) to interpret natural language instructions and autonomously navigate complex three-dimensional environments by grounding language in visual perception. This survey provides a critical and analytical review of the Aerial VLN field, with particular attention to the recent integration of large language models (LLMs) and vision-language models (VLMs). We first formally introduce the Aerial VLN problem and define two interaction paradigms: single-instruction and dialog-based, as foundational axes. We then organize the body of Aerial VLN methods into a taxonomy of five architectural categories: sequence-to-sequence and attention-based methods, end-to-end LLM/VLM methods, hierarchical methods, multi-agent methods, and dialog-based navigation methods. For each category, we systematically analyze design rationales, technical trade-offs, and reported performance. We critically assess the evaluation infrastructure for Aerial VLN, including datasets, simulation platforms, and metrics, and identify their gaps in scale, environmental diversity, real-world grounding, and metric coverage. We consolidate cross-method comparisons on shared benchmarks and analyze key architectural trade-offs, including discrete versus continuous actions, end-to-end versus hierarchical designs, and the simulation-to-reality gap. Finally, we synthesize seven concrete open problems: long-horizon instruction grounding, viewpoint robustness, scalable spatial representation, continuous 6-DoF action execution, onboard deployment, benchmark standardization, and multi-UAV swarm navigation, with specific research directions grounded in the evidence presented throughout the survey.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。