arXiv:2509.06312eess.SYcs.LG2025-09被引 3

用多模态大模型提升低空无人机意图识别能力

Enhancing Low-Altitude Airspace Security: MLLM-Enabled UAV Intent Recognition

  • 融合无人机实时飞行与载荷数据,结合环境与先验知识推理意图
  • 在低空对抗场景中验证了架构可行性,实现精准意图判断
  • 适合低空安防、无人机管控系统研发人员参考

低空经济快速发展,对非合作无人机的有效感知与意图识别提出迫切需求。多模态大语言模型(MLLM)具备先进生成推理能力,为该任务提供新思路。本文提出一种基于MLLM的无人机意图识别架构:通过多模态感知系统获取无人机实时载荷与运动信息,生成结构化输入;再由MLLM结合环境信息、先验知识与战术偏好,输出意图识别结果。我们回顾相关工作并展示其在该架构下的进展。进一步开展低空对抗场景用例,验证架构可行性,并为实际系统设计提供关键洞察。最后讨论未来挑战,提出应用优化建议。

原文摘要 · Abstract (English)

The rapid development of the low-altitude economy emphasizes the critical need for effective perception and intent recognition of non-cooperative unmanned aerial vehicles (UAVs). The advanced generative reasoning capabilities of multimodal large language models (MLLMs) present a promising approach in such tasks. In this paper, we focus on the combination of UAV intent recognition and the MLLMs. Specifically, we first present an MLLM-enabled UAV intent recognition architecture, where the multimodal perception system is utilized to obtain real-time payload and motion information of UAVs, generating structured input information, and MLLM outputs intent recognition results by incorporating environmental information, prior knowledge, and tactical preferences. Subsequently, we review the related work and demonstrate their progress within the proposed architecture. Then, a use case for low-altitude confrontation is conducted to demonstrate the feasibility of our architecture and offer valuable insights for practical system design. Finally, the future challenges are discussed, followed by corresponding strategic recommendations for further applications.

无人机识别多模态模型低空安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。