arXiv:2504.21432cs.ROcs.CV2025-04被引 27

让无人机听懂人话,自动规划飞行路径

UAV-VLN: End-to-End Vision Language guided Navigation for UAVs

  • 用大语言模型理解自然语言指令并融合视觉感知
  • 在新环境和新指令下准确率显著提升,轨迹更高效
  • 适合需要人机交互的无人机自主导航场景

人工智能自主导航的核心挑战是如何让智能体基于自然语言指令,在未见过的环境中实现真实有效的导航。本文提出UAV-VLN,一种面向无人机的端到端视觉-语言导航框架,通过将大语言模型(LLMs)与视觉感知无缝结合,实现人机交互式导航。系统可解析自由形式的自然语言指令,将其与视觉观测对齐,并在多样环境中规划可行的空中路径。UAV-VLN利用大语言模型的常识推理能力解析高层语义目标,同时通过视觉模型检测并定位环境中的语义相关物体。通过多模态融合,无人机能够推理空间关系、消解指令歧义,并在极少任务特定监督下实现上下文感知行为。为保障决策鲁棒性与可解释性,框架引入跨模态对齐机制,使语言意图与视觉上下文精确匹配。我们在多种室内外导航场景中评估该系统,结果表明其在新指令与新环境上具有强泛化能力,指令遵循准确率与轨迹效率均显著提升,验证了大语言模型驱动的视觉-语言接口在安全、直观、通用无人机自主中的潜力。

原文摘要 · Abstract (English)

A core challenge in AI-guided autonomy is enabling agents to navigate realistically and effectively in previously unseen environments based on natural language commands. We propose UAV-VLN, a novel end-to-end Vision-Language Navigation (VLN) framework for Unmanned Aerial Vehicles (UAVs) that seamlessly integrates Large Language Models (LLMs) with visual perception to facilitate human-interactive navigation. Our system interprets free-form natural language instructions, grounds them into visual observations, and plans feasible aerial trajectories in diverse environments. UAV-VLN leverages the common-sense reasoning capabilities of LLMs to parse high-level semantic goals, while a vision model detects and localizes semantically relevant objects in the environment. By fusing these modalities, the UAV can reason about spatial relationships, disambiguate references in human instructions, and plan context-aware behaviors with minimal task-specific supervision. To ensure robust and interpretable decision-making, the framework includes a cross-modal grounding mechanism that aligns linguistic intent with visual context. We evaluate UAV-VLN across diverse indoor and outdoor navigation scenarios, demonstrating its ability to generalize to novel instructions and environments with minimal task-specific training. Our results show significant improvements in instruction-following accuracy and trajectory efficiency, highlighting the potential of LLM-driven vision-language interfaces for safe, intuitive, and generalizable UAV autonomy.

无人机导航视觉语言大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。