arXiv:2509.13352cs.AIcs.RO2025-09被引 10

用大模型让无人机自主决策,能查资料、会推理、懂协作。

Agentic UAVs: LLM-Driven Autonomy with Integrated Tool-Calling and Cognitive Reasoning

论文配图:Agentic UAVs: LLM-Driven Autonomy with Integrated Tool-Calling and Cognitive Reasoning
图 1 · 摘自论文原文
  • 用大模型驱动无人机推理,支持查数据库和调外部工具。
  • 搜救模拟中检测率提升至91%,行动建议正确率92%。
  • 适合需要智能决策的救援、安防等复杂任务场景。

无人机在国防、监控和灾害响应中应用日益广泛,但多数系统仍处于SAE Level 2至3自主水平。其依赖规则控制与专用人工智能,难以适应动态不确定任务。现有架构缺乏上下文感知推理、自主决策能力,也未整合外部系统。尤其缺乏利用具备工具调用功能的大语言模型(LLM)进行实时知识获取。本文提出Agentic UAVs框架,包含感知、推理、行动、集成和学习五层结构。该框架通过大模型驱动的推理、数据库查询及与第三方系统交互,提升无人机自主性。原型基于ROS 2与Gazebo构建,采用YOLOv11进行目标检测,GPT-4负责推理,本地部署Gemma 3模型。在模拟搜救任务中,智能无人机检测置信度达0.79(对比0.72),人员检测率达91%(对比75%),正确行动建议比例达92%(对比4.5%)。结果表明,适度计算开销可显著提升自主等级与系统集成能力。

原文摘要 · Abstract (English)

Unmanned Aerial Vehicles (UAVs) are increasingly used in defense, surveillance, and disaster response, yet most systems still operate at SAE Level 2 to 3 autonomy. Their dependence on rule-based control and narrow AI limits adaptability in dynamic and uncertain missions. Current UAV architectures lack context-aware reasoning, autonomous decision-making, and integration with external systems. Importantly, none make use of Large Language Model (LLM) agents with tool-calling for real-time knowledge access. This paper introduces the Agentic UAVs framework, a five-layer architecture consisting of Perception, Reasoning, Action, Integration, and Learning. The framework enhances UAV autonomy through LLM-driven reasoning, database querying, and interaction with third-party systems. A prototype built with ROS 2 and Gazebo combines YOLOv11 for object detection with GPT-4 for reasoning and a locally deployed Gemma 3 model. In simulated search-and-rescue scenarios, agentic UAVs achieved higher detection confidence (0.79 compared to 0.72), improved person detection rates (91% compared to 75%), and a major increase in correct action recommendations (92% compared to 4.5%). These results show that modest computational overhead can enable significantly higher levels of autonomy and system-level integration.

无人机大模型自主决策工具调用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。