arXiv:2606.31073cs.AIcs.MA2026-06

为多无人机协作任务规划打造专用大模型评测平台,解决真实飞行约束下的智能决策难题。

MultiUAV-Plat: An LLM-Oriented Platform, Benchmark and Framework for Multi-UAV Collaborative Task Planning

论文配图:MultiUAV-Plat: An LLM-Oriented Platform, Benchmark and Framework for Multi-UAV Collaborative Task Planning
图 1 · 摘自论文原文
  • 构建轻量级仿真平台,支持大模型通过真实工具交互完成任务
  • 涵盖75个任务场景、1500条自然语言指令,任务通过率达57.9%
  • 适合研究大模型驱动的无人机群协同与自主规划的学者和开发者

大型语言模型(LLMs)为高级机器人任务规划提供了有前景的接口,但其在多无人机协作中的应用仍缺乏系统评估。现有无人机模拟器多聚焦动力学、感知或低层控制,而现有大模型智能体基准又难以体现空中机器人的约束,如部分可观测性、空间覆盖、无人机分配及多机协调。为此,我们提出MultiUAV-Plat,一个面向大模型智能体的轻量级、易用型多无人机协作任务规划仿真平台。该平台提供简洁的RESTful API、面向智能体的观测信息、基于角色的信息访问、隐藏的验证逻辑及可选2D/3D可视化,使智能体通过真实工具交互完成任务,而非依赖特权模拟器访问。基于此平台,我们构建了MultiUAV-Plat Benchmark,包含75个任务会话、1500条自然语言任务和9396次验证检查,覆盖目标分配、区域搜索、区域分配与巡逻等场景。我们进一步提出Agent4Drone框架,将多无人机行为结构化为记忆、观测、任务理解、规划、执行与验证。在完整配对基准对比中,Agent4Drone达到57.9%的任务通过率、74.6%的平均任务验证通过率和72.0%的全局验证通过率,显著优于ReAct基线的30.6%、47.9%和43.1%。同时,总失败任务率从32.4%降至12.9%。结果表明,MultiUAV-Plat与MultiUAV-Plat Benchmark为在真实信息与执行约束下研究大模型驱动的多无人机自主性提供了可复现的基础。

原文摘要 · Abstract (English)

Large language models (LLMs) provide a promising interface for high-level robotic task planning, but their use in multi-UAV collaboration remains difficult to evaluate systematically. Existing UAV simulators mainly emphasize dynamics, perception, or low-level control, while existing LLM-agent benchmarks rarely capture aerial-robotics constraints such as partial observability, spatial coverage, UAV assignment, and multi-vehicle coordination. To bridge this gap, we present MultiUAV-Plat, a lightweight, easy-to-use, LLM-agent-oriented simulation platform for multi-UAV collaborative task planning. The platform exposes concise RESTful APIs, agent-facing observations, role-based information access, hidden validation logic, and optional 2D/3D visualization, allowing agents to solve missions through realistic tool interaction rather than privileged simulator access. Built on this platform, the MultiUAV-Plat Benchmark contains 75 mission sessions, 1500 natural-language tasks, and 9396 validation checks across target assignment, area search, and area assignment and patrol scenarios. We further propose Agent4Drone, a task-specific LLM agent framework that structures multi-UAV behavior into memory, observation, task understanding, planning, execution, and verification. In a full paired benchmark comparison, Agent4Drone achieves a 57.9% task pass rate, a 74.6% average task check pass rate, and a 72.0% global check pass rate, substantially outperforming a ReAct baseline at 30.6%, 47.9%, and 43.1%, respectively. Agent4Drone also reduces the total failed task rate from 32.4% to 12.9%. These results demonstrate that MultiUAV-Plat and MultiUAV-Plat Benchmark provide a reproducible foundation for studying LLM-driven multi-UAV autonomy under realistic information and execution constraints.

无人机群大模型任务规划仿真平台

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。