arXiv:2606.02313cs.RO2026-06

用专家数据提升视觉语言模型在无人机导航中的意图对齐能力

Towards Precise Intent-Aligned VLA Aerial Navigation via Expert-Guided GRPO

论文配图:Towards Precise Intent-Aligned VLA Aerial Navigation via Expert-Guided GRPO
图 1 · 摘自论文原文
  • 引入专家引导的分组相对策略优化,结合少量专家数据增强在线训练
  • 在复杂指令下成功率提升2.13倍,意图对齐性能提高60.9%
  • 适合需要高精度指令理解的无人机自主导航场景

视觉-语言-动作(VLA)模型为无人飞行器通过细粒度指令完成复杂任务提供了端到端的前景。然而,标准监督微调存在数据稀缺、泛化能力弱及对细微人类意图监督不足的问题。强化学习微调可通过可设计反馈缓解这些问题,但在空域广阔的连续空间中探索效率低。为此,我们提出一种高效的基于VLA的无人机导航强化学习框架。核心是EG-GRPO(专家引导的组相对策略优化),通过少量专家数据增强在线采样。同时设计异构流水线,实现仿真与推理并行,将采样时间减少43.5%。在多种由复杂人类意图定义的任务中,该方法使成功率提升至监督微调基线的2.13倍,意图对齐性能提高60.9%。结果表明,该框架可推动无人机导航向精准意图对齐迈进。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models offer a promising end-to-end paradigm for unmanned aerial vehicles (UAVs) to accomplish complex tasks specified by fine-grained instructions. However, standard supervised fine-tuning (SFT) suffers from data scarcity, limited generalization, and weak supervision for nuanced and complicated human intents. Reinforcement fine-tuning offers a natural way to mitigate these challenges and align policy behaviors with human intents through designable feedback, but applying it to aerial navigation remains challenging due to inefficient exploration in expansive continuous spaces. To address these challenges, we introduce an efficient reinforcement learning (RL) framework for VLA-based aerial navigation. At its core, we propose EG-GRPO (Expert-Guided Group Relative Policy Optimization) to augment online rollouts with few-shot expert data. Additionally, we design a heterogeneous pipeline enabling parallel simulation and inference, which reduces rollout time by 43.5%. Across multiple tasks specified by complex human intents, EG-GRPO improves the success rate to 2.13x that of the SFT baseline, while improving intent alignment performance by 60.9%. These results demonstrate that our framework can move aerial navigation toward precise intent-aligned flight.

无人机导航视觉语言动作强化学习意图对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。