arXiv:2505.12835cs.CLcs.CV2025-05EMNLP被引 30

用视觉语言模型提升无人机导航的通用性与可解释性

FlightGPT: Towards Generalizable and Interpretable UAV Vision-and-Language Navigation with Vision-Language Models

  • 基于视觉语言模型分两阶段训练,先微调再优化策略
  • 在未见环境中成功率达91.2%,比最强基线高9.22%
  • 引入思维链机制,让决策过程更清晰可读

无人机视觉语言导航在灾害响应、物流配送和城市巡检中至关重要。现有方法普遍存在多模态融合不足、泛化能力弱和可解释性差的问题。为此,我们提出基于视觉语言模型的FlightGPT框架,采用两阶段训练:首先使用高质量示范进行监督微调以优化初始化和结构化推理;随后通过组相对策略优化(GRPO)算法,在综合考虑目标达成率、推理质量与格式合规性的奖励引导下,提升泛化与适应能力。此外,FlightGPT引入基于思维链(CoT)的推理机制,增强决策可解释性。在城市尺度数据集CityNav上的大量实验表明,FlightGPT在所有场景下均达到最先进性能,未见环境中的成功率比最强基线高出9.22%。代码已公开。

原文摘要 · Abstract (English)

Unmanned Aerial Vehicle (UAV) Vision-and-Language Navigation (VLN) is vital for applications such as disaster response, logistics delivery, and urban inspection. However, existing methods often struggle with insufficient multimodal fusion, weak generalization, and poor interpretability. To address these challenges, we propose FlightGPT, a novel UAV VLN framework built upon Vision-Language Models (VLMs) with powerful multimodal perception capabilities. We design a two-stage training pipeline: first, Supervised Fine-Tuning (SFT) using high-quality demonstrations to improve initialization and structured reasoning; then, Group Relative Policy Optimization (GRPO) algorithm, guided by a composite reward that considers goal accuracy, reasoning quality, and format compliance, to enhance generalization and adaptability. Furthermore, FlightGPT introduces a Chain-of-Thought (CoT)-based reasoning mechanism to improve decision interpretability. Extensive experiments on the city-scale dataset CityNav demonstrate that FlightGPT achieves state-of-the-art performance across all scenarios, with a 9.22\% higher success rate than the strongest baseline in unseen environments. Our implementation is publicly available.

无人机导航视觉语言模型可解释性强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。