arXiv:2603.28561cs.ROcs.AI2026-03被引 1

用微调大模型实现无人机协同避撞,提升安全与效率。

Fine-Tuning Large Language Models for Cooperative Tactical Deconfliction of Small Unmanned Aerial Systems

  • 基于蓝天空仿真生成规则一致数据,微调大模型决策逻辑。
  • 监督微调使避撞准确率显著提升,近地碰撞大幅减少。
  • 适合关注无人机群智能决策与安全控制的研究者。

小型无人机在低空空域的广泛应用加剧了安全关键环境下可靠战术避撞的需求。战术避撞需在密集、部分可观测、异构多智能体环境中进行短时决策,同时保障协作分离与运行效率。尽管大语言模型具备强大推理能力,但其直接应用于空管仍受限于领域知识不足和输出不一致。本文研究通过微调策略将大模型作为协作多智能体战术避撞决策者,使其输出对齐人类操作员经验。我们基于BlueSky空管仿真器构建了模拟到语言的数据生成流程,生成符合安全规范的避撞数据集。采用预训练Qwen-Math-7B模型,结合低秩适配(LoRA)的监督微调和融合组相对策略优化(GRPO)的偏好微调两种参数高效方法。验证数据集与闭环仿真结果表明,监督微调显著提升了决策准确性、一致性及分离性能,近地空中碰撞显著减少;而GRPO虽增强协作性,但在与异构智能体策略交互时鲁棒性下降。

原文摘要 · Abstract (English)

The growing deployment of small Unmanned Aerial Systems (sUASs) in low-altitude airspaces has increased the need for reliable tactical deconfliction under safety-critical constraints. Tactical deconfliction involves short-horizon decision-making in dense, partially observable, and heterogeneous multi-agent environments, where both cooperative separation assurance and operational efficiency must be maintained. While Large Language Models (LLMs) exhibit strong reasoning capabilities, their direct application to air traffic control remains limited by insufficient domain grounding and unpredictable output inconsistency. This paper investigates LLMs as decision-makers in cooperative multi-agent tactical deconfliction using fine-tuning strategies that align model outputs to human operator heuristics. We propose a simulation-to-language data generation pipeline based on the BlueSky air traffic simulator that produces rule-consistent deconfliction datasets reflecting established safety practices. A pretrained Qwen-Math-7B model is fine-tuned using two parameter-efficient strategies: supervised fine-tuning with Low-Rank Adaptation (LoRA) and preference-based fine-tuning combining LoRA with Group-Relative Policy Optimization (GRPO). Experimental results on validation datasets and closed-loop simulations demonstrate that supervised LoRA fine-tuning substantially improves decision accuracy, consistency, and separation performance compared to the pretrained LLM, with significant reductions in near mid-air collisions. GRPO provides additional coordination benefits but exhibits reduced robustness when interacting with heterogeneous agent policies.

无人机大模型避撞决策强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。