用强化学习自动调优四旋翼飞控参数,飞行中实时优化轨迹跟踪效果。
Reinforcement Learning Based Prediction of PID Controller Gains for Quadrotor UAVs
- 基于DDPG算法的强化学习在线调节PID参数
- 飞行中姿态误差最小,显著优于人工调参
- 适合需要高精度轨迹控制的无人机应用
提出一种基于强化学习(RL)的方法,用于四旋翼无人机姿态控制器的在线参数微调,以提升轨迹跟踪的准确性和有效性。强化学习智能体先在离线环境下通过四旋翼PID姿态控制器训练,随后在仿真和实际飞行中验证。采用深度确定性策略梯度(DDPG)算法,一种离线策略的演员-评论家方法。训练与仿真在Matlab/Simulink及PX4自动驾驶器支持包环境中完成。对比手调参数与强化学习调参的性能表现,结果表明:强化学习在飞行过程中动态调整控制参数,实现最小姿态误差,显著改善了姿态跟踪性能。
原文摘要 · Abstract (English)
A reinforcement learning (RL) based methodology is proposed and implemented for online fine-tuning of PID controller gains, thus, improving quadrotor effective and accurate trajectory tracking. The RL agent is first trained offline on a quadrotor PID attitude controller and then validated through simulations and experimental flights. RL exploits a Deep Deterministic Policy Gradient (DDPG) algorithm, which is an off-policy actor-critic method. Training and simulation studies are performed using Matlab/Simulink and the UAV Toolbox Support Package for PX4 Autopilots. Performance evaluation and comparison studies are performed between the hand-tuned and RL-based tuned approaches. The results show that the controller parameters based on RL are adjusted during flights, achieving the smallest attitude errors, thus significantly improving attitude tracking performance compared to the hand-tuned approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。