arXiv:2604.13654cs.RO2026-04被引 4

无人机视觉语言导航的全景综述,指明技术演进与落地挑战。

Vision-and-Language Navigation for UAVs: Progress, Challenges, and a Research Roadmap

论文配图:Vision-and-Language Navigation for UAVs: Progress, Challenges, and a Research Roadmap
图 1 · 摘自论文原文
  • 按方法演进梳理从模块化到大模型驱动的路径
  • 指出仿真到现实、动态环境感知等核心瓶颈
  • 适合关注无人机智能决策与多智能体协同的研究者

无人飞行器视觉语言导航(UAV-VLN)是具身人工智能中的关键挑战,旨在使无人机理解高层人类指令并在复杂三维环境中执行长时程任务。本文系统性综述该领域,从任务定义到当前技术水平。构建方法学分类体系,追踪技术演进:从早期模块化与深度学习方法,发展至由大基础模型驱动的代理系统,包括视觉-语言模型(VLM)、视觉-语言-动作模型(VLA),以及生成世界模型与VLA架构融合带来的物理感知推理。系统梳理推动研究的关键资源生态:模拟器、数据集与评估指标。通过批判性分析,揭示制约实际部署的主要挑战:仿真到现实的差距、动态户外环境下的鲁棒感知、语言歧义的推理能力,以及大模型在资源受限硬件上的高效部署。基于现有基准与局限,提出前瞻性研究路线图,指引未来在多智能体编队协同与空地协作机器人等前沿方向的探索。

原文摘要 · Abstract (English)

Vision-and-Language Navigation for Unmanned Aerial Vehicles (UAV-VLN) represents a pivotal challenge in embodied artificial intelligence, focused on enabling UAVs to interpret high-level human commands and execute long-horizon tasks in complex 3D environments. This paper provides a comprehensive and structured survey of the field, from its formal task definition to the current state of the art. We establish a methodological taxonomy that charts the technological evolution from early modular and deep learning approaches to contemporary agentic systems driven by large foundation models, including Vision-Language Models (VLMs), Vision-Language-Action (VLA) models, and the emerging integration of generative world models with VLA architectures for physically-grounded reasoning. The survey systematically reviews the ecosystem of essential resources simulators, datasets, and evaluation metrics that facilitates standardized research. Furthermore, we conduct a critical analysis of the primary challenges impeding real-world deployment: the simulation-to-reality gap, robust perception in dynamic outdoor settings, reasoning with linguistic ambiguity, and the efficient deployment of large models on resource-constrained hardware. By synthesizing current benchmarks and limitations, this survey concludes by proposing a forward-looking research roadmap to guide future inquiry into key frontiers such as multi-agent swarm coordination and air-ground collaborative robotics.

无人机导航视觉语言大模型综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。