无人机搜救系统通过视觉语言模型实现快速精准响应
UAV-VLRR: Vision-Language Informed NMPC for Rapid Response in UAV Search and Rescue
- 融合视觉语言模型与大语言模型解析场景并生成指令
- 非线性模型预测控制使无人机响应速度提升33.75%以上
- 适合应急搜救、无人系统指挥等需要快速决策的场景
紧急搜救(SAR)任务常需在复杂环境中快速精准定位目标,传统人工遥控无人机效率低下。为此,本文提出UAV-VLRR(视觉-语言-快速响应)无人机搜救系统。该系统包含两部分:1)多模态系统,结合视觉语言模型(VLM)与ChatGPT-4o(LLM)的自然语言处理能力,实现场景理解;2)内置避障的非线性模型预测控制(NMPC),使无人机根据多模态输出快速安全飞行。实验表明,该方法相较市售自动驾驶仪平均提速33.75%,相比人类飞行员提速54.6%。视频演示:https://youtu.be/KJqQGKKt1xY
原文摘要 · Abstract (English)
Emergency search and rescue (SAR) operations often require rapid and precise target identification in complex environments where traditional manual drone control is inefficient. In order to address these scenarios, a rapid SAR system, UAV-VLRR (Vision-Language-Rapid-Response), is developed in this research. This system consists of two aspects: 1) A multimodal system which harnesses the power of Visual Language Model (VLM) and the natural language processing capabilities of ChatGPT-4o (LLM) for scene interpretation. 2) A non-linearmodel predictive control (NMPC) with built-in obstacle avoidance for rapid response by a drone to fly according to the output of the multimodal system. This work aims at improving response times in emergency SAR operations by providing a more intuitive and natural approach to the operator to plan the SAR mission while allowing the drone to carry out that mission in a rapid and safe manner. When tested, our approach was faster on an average by 33.75% when compared with an off-the-shelf autopilot and 54.6% when compared with a human pilot. Video of UAV-VLRR: https://youtu.be/KJqQGKKt1xY
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。