arXiv:2511.18270cs.RO2025-11被引 1

用物理约束提升大模型在无人机搜索中的决策能力

Skypilot: Fine-Tuning LLM with Physical Grounding for AAV Coverage Search

  • 分两阶段:先生成可行动作,再用蒙特卡洛树搜索优化路径
  • 在2.3万条生成数据上微调Qwen3-4B,推理速度显著提升
  • 真实飞行实验验证了方案高效性,适合智能无人机系统研发

自主飞行器(AAV)在覆盖任务和搜索行动中发挥关键作用。大语言模型(LLM)的进展为增强AAV智能提供了新机遇,可应对区域覆盖优化、动态路径规划与自适应决策等复杂挑战。然而,现有LLM缺乏物理现实约束,导致空间推理与决策中出现幻觉与不可复现问题。为此,我们提出Skypilot——一种基于蒙特卡洛树搜索(MCTS)实现物理接地的两阶段框架。第一阶段引入包含生成、重生成、微调与评估的操作空间,并结合物理感知奖励函数确保轨迹可行性;第二阶段在23,000条MCTS生成样本上微调Qwen3-4B,在保持解质量的同时实现显著推理加速。大量数值仿真与真实飞行实验验证了该方法的效率与优越性。详细信息与实验结果详见https://sky-pilot.top。

原文摘要 · Abstract (English)

Autonomous aerial vehicles (AAVs) have played a pivotal role in coverage operations and search missions. Recent advances in large language models (LLMs) offer promising opportunities to augment AAV intelligence. These advances help address complex challenges like area coverage optimization, dynamic path planning, and adaptive decision-making. However, the absence of physical grounding in LLMs leads to hallucination and reproducibility problems in spatial reasoning and decision-making. To tackle these issues, we present Skypilot, an LLM-enhanced two-stage framework that grounds language models in physical reality by integrating monte carlo tree search (MCTS). In the first stage, we introduce a diversified action space that encompasses generate, regenerate, fine-tune, and evaluate operations, coupled with physics-informed reward functions to ensure trajectory feasibility. In the second stage, we fine-tune Qwen3-4B on 23,000 MCTS-generated samples, achieving substantial inference acceleration while maintaining solution quality. Extensive numerical simulations and real-world flight experiments validate the efficiency and superiority of our proposed approach. Detailed information and experimental results are accessible at https://sky-pilot.top.

无人机搜索大模型物理约束路径规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。