arXiv:2602.08653cs.RO2026-02被引 2

用安全约束强化学习,实现高速无人机避障飞行

High-Speed Vision-Based Flight in Clutter with Safety-Shielded Reinforcement Learning

  • 融合物理先验的端到端强化学习框架
  • 实测最高速度达7.5米/秒,无碰撞通过密集障碍物
  • 适合需要高动态、强安全性的无人机自主飞行场景

四旋翼无人机在复杂任务中对自主导航和避障能力要求越来越高。传统模块化系统存在累积延迟,而纯强化学习方法通常缺乏形式化安全保证。为此,我们提出一种结合模型驱动安全机制的端到端强化学习框架。训练阶段引入物理感知奖励结构,提供全局导航引导;部署阶段集成实时安全过滤器,将策略输出投影至可证明安全的集合,严格满足避障约束。该混合架构实现了高速飞行与可靠安全的统一。基准测试表明,本方法优于传统规划器及基于可微物理的最新端到端避障方法。大量实验验证了强泛化能力,可在密集障碍物环境及复杂室外森林中实现最高7.5米/秒的可靠高速飞行。

原文摘要 · Abstract (English)

Quadrotor unmanned aerial vehicles (UAVs) are increasingly deployed in complex missions that demand reliable autonomous navigation and robust obstacle avoidance. However, traditional modular pipelines often incur cumulative latency, whereas purely reinforcement learning (RL) approaches typically provide limited formal safety guarantees. To bridge this gap, we propose an end-to-end RL framework augmented with model-based safety mechanisms. We incorporate physical priors in both training and deployment. During training, we design a physics-informed reward structure that provides global navigational guidance. During deployment, we integrate a real-time safety filter that projects the policy outputs onto a provably safe set to enforce strict collision-avoidance constraints. This hybrid architecture reconciles high-speed flight with robust safety assurances. Benchmark evaluations demonstrate that our method outperforms both traditional planners and recent end-to-end obstacle avoidance approaches based on differentiable physics. Extensive experiments demonstrate strong generalization, enabling reliable high-speed navigation in dense clutter and challenging outdoor forest environments at velocities up to 7.5 m/s}.

无人机强化学习安全控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。