arXiv:2509.07707cs.ROcs.SY2025-09被引 3

用强化学习让四旋翼在单桨失效时仍能稳住飞行,提升安全性和鲁棒性。

Fault Tolerant Control of a Quadcopter using Reinforcement Learning

  • 采用动态规划与深度确定性策略梯度结合,解决大状态空间下的控制难题。
  • 在多种初始条件下模拟验证,飞行器在桨叶故障后仍能恢复稳定姿态。
  • 适合需要高可靠性的无人机任务,如救援、巡检等关键应用。

本研究提出一种基于强化学习(RL)的新型控制框架,旨在提升四旋翼在飞行中单桨叶失效情况下的安全性与鲁棒性。针对物理应用场景中维持期望高度以保护硬件和载荷的关键需求,本文探讨了两种RL方法——模型依赖的动态规划(DP)与模型无关的深度确定性策略梯度(DDPG),以应对旋翼失效带来的挑战。尽管DP具有收敛保证但计算开销大,而DDPG虽计算快速但解持续时间受限。通过对现有算法进行改进,控制器被训练以适应大规模连续状态与动作空间,并能在飞行中桨叶失效后恢复至目标状态。为验证所提框架的鲁棒性,在MATLAB环境中对多种初始条件进行了大量仿真,结果表明该方法适用于任务关键型四旋翼系统。进一步对比分析了两种算法在故障飞行系统中的适用性。

原文摘要 · Abstract (English)

This study presents a novel reinforcement learning (RL)-based control framework aimed at enhancing the safety and robustness of the quadcopter, with a specific focus on resilience to in-flight one propeller failure. Addressing the critical need of a robust control strategy for maintaining a desired altitude for the quadcopter to safe the hardware and the payload in physical applications. The proposed framework investigates two RL methodologies Dynamic Programming (DP) and Deep Deterministic Policy Gradient (DDPG), to overcome the challenges posed by the rotor failure mechanism of the quadcopter. DP, a model-based approach, is leveraged for its convergence guarantees, despite high computational demands, whereas DDPG, a model-free technique, facilitates rapid computation but with constraints on solution duration. The research challenge arises from training RL algorithms on large dimensions and action domains. With modifications to the existing DP and DDPG algorithms, the controllers were trained not only to cater for large continuous state and action domain and also achieve a desired state after an inflight propeller failure. To verify the robustness of the proposed control framework, extensive simulations were conducted in a MATLAB environment across various initial conditions and underscoring its viability for mission-critical quadcopter applications. A comparative analysis was performed between both RL algorithms and their potential for applications in faulty aerial systems.

强化学习四旋翼容错控制无人机

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。