让无人机导航更可靠:通过失败感知强化学习提升飞行精度
RecoverFly: A Failure-Aware Reinforcement Learning Post-Training Framework for Aerial Vision-Language Navigation

- 用逐标记强化学习优化语言引导的飞行指令生成
- 重访失败案例,使样本利用率提升30%以上
- 适合需要高鲁棒性的无人机自主导航场景
无人飞行器视觉语言导航(UAV-VLN)要求智能体将视觉观测与语言指令转化为复杂环境中的可靠飞行动作。尽管近期端到端的无人机视觉-语言-动作(UAV-VLA)策略减少了对独立感知、规划和控制模块的依赖,但其行为克隆目标在交互式闭环执行中提供的纠正监督有限。强化学习(RL)虽具潜力,但受限于样本利用效率低、场景分布长尾化及优化过程中的策略分布偏移。为此,我们提出RecoverFly,一种面向端到端UAV-VLA策略的失败感知强化学习后训练框架。具体而言,RecoverFly采用标记级强化学习以稳定优化语法约束的自回归飞行动作,重访未解决的失败案例以增强纠正学习与样本利用,并结合两阶段长尾场景课程与参考策略正则化,在提升场景适应性的同时保留已有能力。在TravelUAV基准上的实验表明,RecoverFly在已见、未见地图和未见物体分割上均取得最优性能。相比AerialVLA初始化,当总回放预算约为训练集规模的30%时,成功率提升3.12至8.37个百分点,验证了其有效性、鲁棒性与泛化能力。
原文摘要 · Abstract (English)
Unmanned aerial vehicle vision-language navigation (UAV-VLN) requires agents to translate visual observations and language instructions into reliable flight actions in complex environments. Although recent end-to-end UAV vision-language-action (UAV-VLA) policies reduce reliance on separately designed perception, planning, and control modules, their behavior-cloning objectives provide limited corrective supervision for interactive closed-loop execution. Reinforcement learning (RL) offers a promising solution, while its effectiveness is constrained by inefficient use of samples, long-tailed scene distributions, and policy distribution shift during optimization. To this end, we propose RecoverFly, a failure-aware RL post-training framework for end-to-end UAV-VLA policies. Specifically, RecoverFly adapts token-level RL for stable optimization of grammar-constrained autoregressive UAV actions, revisits unresolved failure cases to strengthen corrective learning and sample utilization, and combines a two-stage long-tail scene curriculum with reference-policy regularization to improve scene adaptation while preserving acquired capabilities. Experiments on the TravelUAV benchmark demonstrate that RecoverFly achieves the best performance on the seen, unseen-map, and unseen-object splits. Moreover, compared to the AerialVLA initialization, RecoverFly improves success rate by 3.12 to 8.37 percentage points under a total rollout budget of about 30\% of the training-set size, validating its effectiveness, robustness, and generalization capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。