解决飞行控制中在线学习导致的振荡问题,提升跟踪精度与鲁棒性。
Temporally smoothed incremental model-based heuristic dynamic programming for command-filtered cascaded online learning flight control
- 在IHDP框架中加入时间尺度平滑正则化,抑制控制振荡
- 用低通滤波器降低高频俯仰率指令,减少执行器负担
- 自适应调节平滑惩罚权重,实现稳定控制与快速学习平衡
近似动态规划(ADP)通过自适应评价方法实现控制律的在线自适应。然而,在非线性动力学下,其应用于在线学习攻角(AoA)跟踪控制时仍面临控制动作振荡的问题,可能引发系统振荡、降低跟踪性能并增加执行器负担。为此,本文提出两项创新:(1)将时间尺度策略平滑性正则化引入增量式模型-启发式动态规划(IHDP)框架;(2)采用低通滤波器衰减高频俯仰率指令。此外,设计了原始-对偶方法,根据预设平滑性准则自适应调整策略目标中的平滑惩罚权重。跟踪控制仿真表明,所提方法有效降低控制系统振荡,提升跟踪性能,并在训练后表现出对模型不确定性的鲁棒性。
原文摘要 · Abstract (English)
Approximate Dynamic Programming (ADP) enables online adaptation of control laws through adaptive critic methods. However, its application to online learning Angle-of-Attack (AoA) tracking control remains challenging due to oscillatory control actions under nonlinear dynamics, which may induce system oscillations, degrade tracking performance, and increase actuator effort. To address this challenge, this paper proposes two innovations for online learning flight control: (1) incorporating temporal-scale policy smoothness regularization into the Incremental Model-based Heuristic Dynamic Programming (IHDP) framework; and (2) employing a low-pass filter to attenuate high-frequency pitch-rate commands. Furthermore, a primal-dual approach is developed to adaptively adjust the smoothness penalty weight in the policy objective according to a prescribed smoothness criterion. Tracking control simulations demonstrate that the proposed methods reduce control system oscillations, improve tracking performance, and exhibit post-training policy robustness to model uncertainties.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。