用大模型让无人机在复杂环境自主执行开放任务
General-Purpose Aerial Intelligent Agents Empowered by Large Language Models

- 将大模型推理与飞行控制深度结合,实现边端部署的实时决策
- 140亿参数模型在220瓦功耗下实现每秒5-6个词的推理速度
- 适合需要自主规划与动态响应的野外巡检、探索等场景
大型语言模型(LLM)为无人飞行器(UAV)带来新可能,但现有系统仍受限于软硬件协同设计难题,仅能执行预设任务。本文提出首个能够通过紧密集成大模型推理与机器人自主性的空中智能代理,实现开放世界任务执行。所提出的软硬件协同系统解决了两大核心问题:(1) 基于边缘优化计算平台实现机载运行,使140亿参数模型在220瓦峰值功耗下达到5-6个词/秒的推理速度;(2) 采用双向认知架构,融合慢速深思型规划(基于LLM的任务规划)与快速反应式控制(状态估计、建图、避障及运动规划)。初步原型验证表明,该系统在通信受限环境中具备可靠的任务规划与场景理解能力,适用于甘蔗监测、电网巡检、矿洞勘探和生物观测等应用。本工作建立了一种新型具身空中人工智能框架,弥合了开放环境中任务规划与机器人自主性的鸿沟。
原文摘要 · Abstract (English)
The emergence of large language models (LLMs) opens new frontiers for unmanned aerial vehicle (UAVs), yet existing systems remain confined to predefined tasks due to hardware-software co-design challenges. This paper presents the first aerial intelligent agent capable of open-world task execution through tight integration of LLM-based reasoning and robotic autonomy. Our hardware-software co-designed system addresses two fundamental limitations: (1) Onboard LLM operation via an edge-optimized computing platform, achieving 5-6 tokens/sec inference for 14B-parameter models at 220W peak power; (2) A bidirectional cognitive architecture that synergizes slow deliberative planning (LLM task planning) with fast reactive control (state estimation, mapping, obstacle avoidance, and motion planning). Validated through preliminary results using our prototype, the system demonstrates reliable task planning and scene understanding in communication-constrained environments, such as sugarcane monitoring, power grid inspection, mine tunnel exploration, and biological observation applications. This work establishes a novel framework for embodied aerial artificial intelligence, bridging the gap between task planning and robotic autonomy in open environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。