用哈密顿-雅可比方法优化无人机指令监督,提升复杂飞行中的控制精度。
Autopilot-Preserving Residual Q-Learning with HJB-Inspired Finite-Action Risk Filtering for Fixed-Wing UAV Command Supervision

- 在不变的自动驾驶仪上层添加学习型监督器,通过有限动作集调整指令参考值。
- 路径跟踪均方误差降至44.809米,相比基线降低86.77%。
- 结合李雅普诺夫与屏障函数思想,确保安全退避机制始终可用。
固定翼无人机需在风、阵风和湍流下维持空速、高度和航向参考值,各通道耦合紧密,修正一个变量可能恶化另一个。传统自动驾驶仪对飞机姿态稳定良好,但在强侧风下执行急转弯时适应性差;而直接作用于舵面的强化学习策略在执行器接口处集中探索风险。本文将学习型监督器置于不变的自动驾驶仪之上而非内部:它从有限有界动作集中选择一个残差,作用于命令空速、高度和航向;修改后的参考值在进入自动驾驶仪前被投影至可接受命令范围内,自动驾驶仪仍为唯一面向执行器的控制器。关键创新在于残差的选择方式:基于哈密顿-雅可比(HJB)方程精神,使用半离散值迭代批评器计算候选动作的HJB残差,按无操作相对哈密顿优势排序,并通过受控制李雅普诺夫与控制屏障启发的有限动作防护罩过滤,始终保留无操作备用方案。在一个固定12状态运行环境(包含植株、自动驾驶仪和执行器模型)中,该方法将平均均方根路径跟踪误差降至44.809米,优于基线自动驾驶仪的338.617米和表格型Q残差的88.809米,分别降低86.77%和49.54%。性能提升集中在基线表现最差区域,同时伴随空速误差小幅上升,表明无方法在所有指标上全面占优。本文提出这一保持自动驾驶仪不变的残差指令监督设计,并完整报告其权衡关系。
原文摘要 · Abstract (English)
A fixed-wing UAV must hold airspeed, altitude, and heading references under wind, gusts, and turbulence, channels coupled so that correcting one can degrade another. Classical autopilots stabilize the airframe well but adapt poorly when a hard crosswind meets an aggressive turn, while reinforcement-learning (RL) policies acting directly on the surfaces concentrate exploration risk at the actuator interface. We place a learned supervisor above an unchanged autopilot rather than inside it: it selects a residual from a finite, bounded action set on the commanded airspeed, altitude, and heading; the modified reference is projected into an admissible command envelope before reaching the autopilot, which stays the only actuator-facing controller. What is new is how the residual is chosen. HJB residual scores candidates with a semi-discrete value-iteration critic in the spirit of the Hamilton-Jacobi-Bellman (HJB) equation, ranks them by a no-op-relative Hamiltonian advantage, and filters them through a control-Lyapunov- and control-barrier-inspired finite-action shield that always keeps a no-op fallback. On a shared 12-state runtime holding the plant, autopilot, and actuator model fixed, so the comparison is at the package level, HJB residual lowers mean RMS path-tracking error to 44.809 m, against 338.617 m for the baseline autopilot and 88.809 m for a tabular-Q residual, an 86.77% reduction over the baseline and 49.54% over Q-learning. The gain concentrates where the baseline fails worst and comes with a measured rise in airspeed error, so no method dominates every metric. We present this autopilot-preserving residual command-supervision design and benchmark with its trade-offs reported intact.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。