arXiv:2607.01794cs.ROcs.AI2026-07

轻量化安全强化学习让无人机在复杂环境中更稳更快导航

Lightweight Safe Reinforcement Learning for End-to-End UAV Navigation

论文配图:Lightweight Safe Reinforcement Learning for End-to-End UAV Navigation
图 1 · 摘自论文原文
  • 用轻量网络融合稀疏感知,生成避障风险特征
  • 结合分层控制与约束优化,飞行成功率超基线18%以上
  • 适合算力受限的机载系统,尤其高速飞行场景

随着自主空载系统快速发展,无人机在巡检、环境监测和救援等场景中广泛应用,对可靠自主导航的需求日益增长。然而,在密集环境下的稀疏感知和动态约束条件下,自主导航仍具挑战性。现有强化学习方法缺乏显式安全机制,导致探索不安全、训练不稳定、高飞速时行为危险。即使采用安全强化学习,也常通过投影策略输出到安全动作集来保证安全,可能引发不稳定性。同时,许多基于学习的方法依赖密集输入或大模型,增加计算负担,限制轻量化机载部署。针对上述问题,本文提出一种融合感知-控制的安全约束框架。采用轻量网络,利用非对称和深度可分离卷积将稀疏观测编码为碰撞风险感知特征。在分层控制架构下,将任务建模为带约束的马尔可夫决策过程,并使用基于拉格朗日的受约束PPO算法求解。课程学习进一步提升训练稳定性。在不同障碍密度和飞行速度下的实验表明,该方法在成功率、安全性与效率方面均优于现有强化学习基线。

原文摘要 · Abstract (English)

With the rapid development of autonomous aerial systems, Unmanned Aerial Vehicles (UAVs) are increasingly deployed in applications such as inspection, environmental monitoring, and rescue, creating growing demand for reliable autonomous navigation. However, autonomous UAV navigation in dense environments remains challenging under sparse perception and dynamic constraints. Most reinforcement learning (RL) methods lack explicit safety mechanisms, leading to unsafe exploration, unstable training, and risky behaviors, especially during high-speed flight. Even in safe RL approaches, safety is often enforced by projecting policy outputs onto a safe action set, which may introduce instability. Meanwhile, many learning-based methods rely on dense inputs or large networks, increasing computational burden and limiting lightweight onboard deployment. Facing the above challenges, we propose a safety-constrained perception-control integrated framework for UAV navigation. A lightweight network encodes sparse observations into collision-risk-aware features using asymmetric and depthwise separable convolutions. We formulate the task as a constrained Markov decision process within a hierarchical control architecture and solve it using a Lagrangian-based safe PPO algorithm. Curriculum learning further improves training stability. Experiments with varying obstacle densities and flight speeds demonstrate higher success rates, improved safety, and better efficiency than existing reinforcement learning baselines.

无人机导航强化学习轻量化安全控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。