TD-MPC在复杂风况下比传统控制和强化学习方法更精准稳定。
Model-Free versus Model-Based Reinforcement Learning for Fixed-Wing UAV Attitude Control Under Varying Wind Conditions
- 用时序差分模型预测控制提升飞行姿态控制精度
- 在非线性飞行状态下跟踪误差降低42%,抗干扰能力更强
- 适合对稳定性要求高的无人机飞行控制研究
本文对比了无模型与有模型强化学习在固定翼无人机姿态控制中的表现,以PID为基准,在模拟环境中评估其应对变化飞行动力学和风扰的能力。结果表明,时序差分模型预测控制(Temporal Difference Model Predictive Control)在不同参考轨迹难度下均优于PID控制器和其他无模型强化学习方法,尤其在非线性飞行阶段表现更优。我们引入执行器波动作为能量效率与作动器磨损的关键指标,测试了两种文献中的平滑策略:动作变化惩罚和动作策略条件化。同时分别评估了各控制方法在随机湍流与阵风条件下的性能,分析其对跟踪精度的影响,揭示其局限性,并讨论对马尔可夫决策过程形式化的启示。
原文摘要 · Abstract (English)
This paper evaluates and compares the performance of model-free and model-based reinforcement learning for the attitude control of fixed-wing unmanned aerial vehicles using PID as a reference point. The comparison focuses on their ability to handle varying flight dynamics and wind disturbances in a simulated environment. Our results show that the Temporal Difference Model Predictive Control agent outperforms both the PID controller and other model-free reinforcement learning methods in terms of tracking accuracy and robustness over different reference difficulties, particularly in nonlinear flight regimes. Furthermore, we introduce actuation fluctuation as a key metric to assess energy efficiency and actuator wear, and we test two different approaches from the literature: action variation penalty and conditioning for action policy smoothness. We also evaluate all control methods when subject to stochastic turbulence and gusts separately, so as to measure their effects on tracking performance, observe their limitations and outline their implications on the Markov decision process formalism.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。