用大模型让无人机群更懂人意,救灾效率提升64%。
An LLM-based Framework for Human-Swarm Teaming Cognition in Disaster Search and Rescue
- 用大模型理解人类口令和绘图指令,自动分解任务
- 任务完成时间缩短64.2%,成功率提升7%
- 适合应急救援、军事协同等高压场景
大规模灾难搜救(SAR)常受复杂地形和通信中断困扰。无人机群虽能执行广域搜索与物资投送,但有效协调给操作员带来巨大认知负担。核心瓶颈在于‘意图-行动鸿沟’——在高强度压力下,将高层救援目标转化为低层群控指令易出错。本文提出一种基于大语言模型(LLM)的CRF系统,通过语音或图形标注捕捉操作员意图,利用LLM作为认知引擎实现意图理解、分层任务分解与群组规划。该闭环框架使无人机群成为主动协作伙伴,实时反馈并减少人工监控需求,显著提升搜救效能。在模拟搜救场景中测试表明,相比传统命令界面,该方法使任务完成时间减少约64.2%,任务成功率提高7%,主观认知负荷降低42.9%(NASA-TLX评分)。本研究验证了大模型在高风险人机协同中的应用潜力。
原文摘要 · Abstract (English)
Large-scale disaster Search And Rescue (SAR) operations are persistently challenged by complex terrain and disrupted communications. While Unmanned Aerial Vehicle (UAV) swarms offer a promising solution for tasks like wide-area search and supply delivery, yet their effective coordination places a significant cognitive burden on human operators. The core human-machine collaboration bottleneck lies in the ``intention-to-action gap'', which is an error-prone process of translating a high-level rescue objective into a low-level swarm command under high intensity and pressure. To bridge this gap, this study proposes a novel LLM-CRF system that leverages Large Language Models (LLMs) to model and augment human-swarm teaming cognition. The proposed framework initially captures the operator's intention through natural and multi-modal interactions with the device via voice or graphical annotations. It then employs the LLM as a cognitive engine to perform intention comprehension, hierarchical task decomposition, and mission planning for the UAV swarm. This closed-loop framework enables the swarm to act as a proactive partner, providing active feedback in real-time while reducing the need for manual monitoring and control, which considerably advances the efficacy of the SAR task. We evaluate the proposed framework in a simulated SAR scenario. Experimental results demonstrate that, compared to traditional order and command-based interfaces, the proposed LLM-driven approach reduced task completion time by approximately $64.2\%$ and improved task success rate by $7\%$. It also leads to a considerable reduction in subjective cognitive workload, with NASA-TLX scores dropping by $42.9\%$. This work establishes the potential of LLMs to create more intuitive and effective human-swarm collaborations in high-stakes scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。