arXiv:2603.07181cs.CV2026-03被引 1

让无人机像人一样思考路线,边看图边理解指令导航

FreeFly-Thinking : Aligning Chain-of-Thought Reasoning with Continuous UAV Navigation

  • 将语言指令转化为分步推理,指导无人机在城市环境中导航
  • 在未见过的测试场景中实现高效稳定导航,表现优于传统黑箱模型
  • 适合研究无人机智能导航、多模态推理与自主决策的学者

视觉-语言导航旨在使智能体理解自然语言指令,并在真实环境中执行恰当的导航动作。现有研究多集中于室内场景,对复杂室外环境关注较少。当前无人机视觉-语言导航模型通常为黑箱式,缺乏显式推理过程。本文提出 FreeFly-thinking,一种端到端的视觉-语言导航框架,受 OpenFly 提出的城市建筑环境启发,将无人机的视角图像与语言指令转化为一系列导航动作。首先构建用于导航任务的无人机数据集,随后引入自然语言链式思维(Chain-of-Thought)推理机制。采用两阶段训练策略:监督微调与强化学习微调。在未见测试集上的实验表明,该方法在无人机导航任务中展现出优异的鲁棒性与效率。

原文摘要 · Abstract (English)

Vision-Language Navigation aims to enable agents to understand natural language instructions and carry out appropriate navigation actions in real-world environments. Most work focuses on indoor settings, with little research in complex outdoor scenes. Current UAV Vision-and-Language Navigation models typically act as black boxes without explicit reasoning. We introduce FreeFly-thinking, an end-to-end VLN framework that converts the UAV agent's egocentric images and language instructions into a series of actions, inspired by environment of urban architecture proposed by OpenFly. We first construct a UAV dataset for navigation task, and then performing natural language chain of thought. We adopt a two-stage training strategy: Supervised fine-tuning and Reinforcement fine-tuning. Experiments on unseen test demonstrate a strong performance, presenting robustness and efficiency in UAV navigation issue.

无人机导航视觉语言链式思维多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。