arXiv:2508.11196cs.CV2025-08中稿 · Frontiers of Compu…被引 3

针对无人机影像设计轻量视觉语言模型,提升空中场景理解能力。

UAV-VL-R1: Generalizing Vision-Language Models via Supervised Fine-Tuning and Multi-Stage GRPO for UAV Visual Reasoning

  • 结合监督微调与多阶段强化学习,提升推理结构化水平。
  • 零样本准确率比基线高48.17%,小模型超越36倍大的基线。
  • 仅需3.9GB内存,支持资源受限无人机实时部署。

近年来视觉语言模型在自然图像任务中表现出强大泛化能力,但在无人机航拍影像上性能常下降,因图像分辨率高、空间语义复杂且需严格实时响应。为应对挑战,我们提出UAV-VL-R1,一种专为航拍视觉推理设计的轻量级视觉语言模型。采用监督微调(SFT)与多阶段强化学习(RL)相结合的训练方法,利用组相对策略优化(GRPO)算法,通过规则引导奖励和组内策略对齐,促进结构化可解释推理。为此构建了高分辨率视觉问答数据集HRVQA-VL,含50,019个标注样本,涵盖物体计数、交通识别、空间场景推断等八类无人机相关任务。实验表明,UAV-VL-R1零样本准确率比Qwen2-VL-2B-Instruct基线高出48.17%,甚至超越其72B参数版本(大36倍)。消融实验显示,虽SFT增强语义对齐但降低数学任务推理多样性,而基于GRPO的强化学习弥补此缺陷,提升逻辑灵活性与推理鲁棒性。此外,该模型在FP16下仅需3.9GB内存,量化至INT8后降至2.5GB,支持在资源受限的无人机平台实时运行。

原文摘要 · Abstract (English)

Recent advances in vision-language models (VLMs) have demonstrated strong generalization in natural image tasks. However, their performance often degrades on unmanned aerial vehicle (UAV)-based aerial imagery, which features high resolution, complex spatial semantics, and strict real-time constraints. These challenges limit the applicability of general-purpose VLMs to structured aerial reasoning tasks. To address these challenges, we propose UAV-VL-R1, a lightweight VLM explicitly designed for aerial visual reasoning. It is trained using a hybrid method that combines supervised fine-tuning (SFT) and multi-stage reinforcement learning (RL). We leverage the group relative policy optimization (GRPO) algorithm to promote structured and interpretable reasoning through rule-guided rewards and intra-group policy alignment. To support model training and evaluation, we introduce a high-resolution visual question answering dataset named HRVQA-VL, which consists of 50,019 annotated samples covering eight UAV-relevant reasoning tasks, including object counting, transportation recognition, and spatial scene inference. Experimental results show that UAV-VL-R1 achieves a 48.17% higher zero-shot accuracy than the Qwen2-VL-2B-Instruct baseline and even outperforms its 72B-scale variant, which is 36x larger, on multiple tasks. Ablation studies reveal that while SFT improves semantic alignment, it may reduce reasoning diversity in mathematical tasks. GRPO-based RL compensates for this limitation by enhancing logical flexibility and the robustness of inference. Additionally, UAV-VL-R1 requires only 3.9GB of memory under FP16 inference and can be quantized to 2.5GB with INT8, supporting real-time deployment on resource-constrained UAV platforms.

视觉语言模型无人机推理轻量化部署强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。