arXiv:2607.03146cs.RO2026-07

用专家示范训练小型视觉语言动作模型,让无人机听懂指令自动飞。

Exp2VLA: Enabling Vision-Language-Action for Drone Navigation from Expert Demonstrations

论文配图:Exp2VLA: Enabling Vision-Language-Action for Drone Navigation from Expert Demonstrations
图 1 · 摘自论文原文
  • 从强化学习或遥控等专家策略中提取行为数据,蒸馏成可微调的训练集
  • 在多目标场景下,模型能理解语义指令并泛化到未见目标组合
  • 适合想快速部署语言控制无人机的开发者,降低系统集成难度

视觉-语言-动作(VLA)模型通过端到端框架直接关联感知、语言与动作,为机器人控制提供新路径。然而,无人机实际应用仍受限于现有方案计算开销大或复杂环境表现不足。本文提出一种实用的专家蒸馏管道(Exp2VLA),将强化学习、遥控等专家策略生成的行为数据转化为训练数据,用于微调紧凑型VLA模型。该方法使现有控制策略可统一迁移至语言引导导航框架,减少手动集成,降低新行为部署门槛。在模拟到模拟及仿真回路设置下的多目标场景实验表明,微调后的模型能处理多样语义指令,并泛化至未见目标组合。本框架展示了专家策略蒸馏如何推动机电系统从专用控制模块向更灵活、可复用的机器人智能演进。

原文摘要 · Abstract (English)

Vision-language-action (VLA) models open a new path toward intuitive robot control by directly linking perception, language, and action in a single end-to-end framework. Yet for UAVs, practical adoption remains difficult because existing solutions are either computationally heavy or insufficiently capable in complex environments. In this work, we propose a practical expert-distillation pipeline (Exp2VLA) for language-conditioned drone navigation. The core idea is to distill expert behavior, obtained from reinforcement learning, teleoperation, or other controllers, into training data that can be used to fine-tune compact VLA models. This allows existing control strategies to be transferred into a unified language-guided navigation model, reducing manual system integration and lowering the barrier for deploying new robot behaviors. Experiments in both sim-to-sim and simulation-in-the-loop settings across multi-object scenes show that the fine-tuned models can handle varied semantic commands and generalize to unseen target compositions. The proposed framework demonstrates how expert-policy distillation can help mechatronic systems move from specialized control modules toward more flexible and reusable robot intelligence.

无人机导航语言控制模型蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。