让无人机在复杂空域中高效听懂指令飞行,数据少也能精准规划路径。
OpenVLN: Open-world Aerial Vision-Language Navigation
- 用规则策略微调视觉语言模型,少量数据下完成导航任务
- 长程规划器通过价值奖励动态生成精确飞行动作,提升路径质量
- 在TravelUAV上性能超越基线4.34%成功率,适合复杂空域自主导航
视觉语言模型(VLM)已在地面视觉语言导航(VLN)中广泛应用。然而,室外空域环境的高复杂性带来数据获取困难,并对无人机(UAV)的长时程轨迹规划提出挑战,使空中VLN面临新难题。为此,我们提出一种数据高效的开放世界空中视觉语言导航框架(OpenVLN),可在有限数据条件下实现语言引导飞行,并增强复杂空域中的长程轨迹规划能力。具体而言,我们重构强化学习框架以优化VLM用于无人机导航任务,通过基于规则的策略在少量训练数据下高效微调。同时引入长程规划器,利用基于价值的奖励动态生成精准飞行动作。在TravelUAV基准上进行充分实验,涵盖多种数据规模与奖励设置。结果表明,相比基线方法,本方法在成功率达4.34%、最优成功率达6.19%、路径长度加权成功率达4.07%方面均有稳定提升,验证了其在复杂空域中长程无人机导航部署的有效性。
原文摘要 · Abstract (English)
Vision-language models (VLMs) have been widely-applied in ground-based vision-language navigation (VLN). However, the vast complexity of outdoor aerial environments compounds data acquisition challenges and imposes long-horizon trajectory planning requirements on Unmanned Aerial Vehicles (UAVs), introducing novel complexities for aerial VLN. To address these challenges, we propose a data-efficient Open-world aerial Vision-Language Navigation (i.e., OpenVLN) framework, which could execute language-guided flight with limited data constraints and enhance long-horizon trajectory planning capabilities in complex aerial environments. Specifically, we reconfigure a reinforcement learning framework to optimize the VLM for UAV navigation tasks, which can efficiently fine-tune VLM by using rule-based policies under limited training data. Concurrently, we introduce a long-horizon planner for trajectory synthesis that dynamically generates precise UAV actions via value-based rewards. To the end, we conduct sufficient navigation experiments on the TravelUAV benchmark with dataset scaling across diverse reward settings. Our method demonstrates consistent performance gains of up to 4.34% in Success Rate, 6.19% in Oracle Success Rate, and 4.07% in Success weighted by Path Length over baseline methods, validating its deployment efficacy for long-horizon UAV navigation in complex aerial environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。