arXiv:2503.01378cs.RO2025-03被引 27

让无人机实时理解指令并自主决策,提升复杂任务成功率。

CognitiveDrone: A VLA Model and Evaluation Benchmark for Real-Time Cognitive Task Solving and Reasoning in UAVs

  • 基于视觉语言动作模型,结合第一视角图像与文本指令生成4D控制指令。
  • 在复杂任务中成功率达77.2%,较竞品提升超30%。
  • 适合研究智能无人机、认知计算与人机协同的开发者与学者。

本文提出CognitiveDrone,一种专为需高级认知能力的复杂无人飞行器(UAV)任务设计的视觉-语言-动作(VLA)模型。该模型在包含超过8,000条模拟飞行轨迹的数据集上训练,覆盖人类识别、符号理解与推理三大核心类别,能根据第一人称视觉输入与文本指令实时生成4D动作命令。为进一步提升复杂场景表现,提出CognitiveDrone-R1,通过引入额外的视觉-语言模型(VLM)推理模块,在高频控制前简化任务指令。利用开源基准CognitiveDroneBench进行实验评估显示:以竞速为导向的RaceVLA模型总体成功率为31.3%,基础CognitiveDrone模型达59.6%,而CognitiveDrone-R1成功率达到77.2%。结果表明,在关键认知任务中性能提升最高达30%,验证了将先进推理能力融入无人机控制系统的效果。本工作贡献包括构建当前最先进的无人机控制VLA模型,并推出首个针对无人机认知任务评估的专用基准。完整代码库可访问cognitivedrone.github.io。

原文摘要 · Abstract (English)

This paper introduces CognitiveDrone, a novel Vision-Language-Action (VLA) model tailored for complex Unmanned Aerial Vehicles (UAVs) tasks that demand advanced cognitive abilities. Trained on a dataset comprising over 8,000 simulated flight trajectories across three key categories-Human Recognition, Symbol Understanding, and Reasoning-the model generates real-time 4D action commands based on first-person visual inputs and textual instructions. To further enhance performance in intricate scenarios, we propose CognitiveDrone-R1, which integrates an additional Vision-Language Model (VLM) reasoning module to simplify task directives prior to high-frequency control. Experimental evaluations using our open-source benchmark, CognitiveDroneBench, reveal that while a racing-oriented model (RaceVLA) achieves an overall success rate of 31.3%, the base CognitiveDrone model reaches 59.6%, and CognitiveDrone-R1 attains a success rate of 77.2%. These results demonstrate improvements of up to 30% in critical cognitive tasks, underscoring the effectiveness of incorporating advanced reasoning capabilities into UAV control systems. Our contributions include the development of a state-of-the-art VLA model for UAV control and the introduction of the first dedicated benchmark for assessing cognitive tasks in drone operations. The complete repository is available at cognitivedrone.github.io

无人机视觉语言认知推理VLA

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。