arXiv:2506.12710cs.RO2025-06被引 31

用多模态大模型提升无人机集群的智能与自适应能力

Multimodal Large Language Models-Enabled UAV Swarm: Towards Efficient and Intelligent Autonomous Aerial Systems

  • 将多模态大模型融入无人机集群,实现跨模态感知与决策
  • 在森林火灾救援中验证了任务规划、火情评估与协同执行能力
  • 适合研究智能无人系统、人机协作与复杂环境自主决策的学者

多模态大语言模型(MLLM)的突破使AI具备跨文本、图像和视频流的统一感知、推理与自然语言交互能力。与此同时,无人机集群越来越多地应用于动态、高安全要求的任务,亟需快速态势理解与自主适应能力。本文探讨将MLLM与无人机集群结合以增强其智能性与适应性的可能方案。首先概述无人机与MLLM的基本架构与功能;接着分析MLLM如何提升目标检测、自主导航与多智能体协同性能,并提出集成方案;随后以森林火灾救援为实际案例,研究人机交互、集群任务规划、火情评估与任务执行;最后讨论该框架面临的挑战与未来方向。实验演示视频可在线观看:https://youtu.be/zwnB9ZSa5A4。

原文摘要 · Abstract (English)

Recent breakthroughs in multimodal large language models (MLLMs) have endowed AI systems with unified perception, reasoning and natural-language interaction across text, image and video streams. Meanwhile, Unmanned Aerial Vehicle (UAV) swarms are increasingly deployed in dynamic, safety-critical missions that demand rapid situational understanding and autonomous adaptation. This paper explores potential solutions for integrating MLLMs with UAV swarms to enhance the intelligence and adaptability across diverse tasks. Specifically, we first outline the fundamental architectures and functions of UAVs and MLLMs. Then, we analyze how MLLMs can enhance the UAV system performance in terms of target detection, autonomous navigation, and multi-agent coordination, while exploring solutions for integrating MLLMs into UAV systems. Next, we propose a practical case study focused on the forest fire fighting. To fully reveal the capabilities of the proposed framework, human-machine interaction, swarm task planning, fire assessment, and task execution are investigated. Finally, we discuss the challenges and future research directions for the MLLMs-enabled UAV swarm. An experiment illustration video could be found online at https://youtu.be/zwnB9ZSa5A4.

无人机集群多模态大模型智能系统自主决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。