用神经网络在线优化飞行控制,让无人机敏捷穿门并快速抗扰。
Learning Agile Gate Traversal via Analytical Optimal Policy Gradient
- 用离线训练的神经网络预测飞行参考姿态和成本权重,动态调优模型预测控制参数。
- 实测峰值加速度达30 m/s²,干扰后0.85秒内恢复,体感速率扰动超1146度/秒。
- 结合可解释性与高样本效率,适合需要精准避障与强鲁棒性的无人机场景。
穿越狭窄门框是评估四旋翼无人机敏捷与精准飞行的标准挑战。传统模块化自主飞行系统需大量设计与参数调优,而端到端强化学习方法常面临样本效率低、可解释性差及未知扰动下抗干扰能力下降的问题。本文提出一种新型混合框架,通过离线训练的神经网络(NN)在线自适应微调模型预测控制(MPC)参数。该神经网络基于门框角点坐标与当前无人机状态,联合预测参考姿态和代价函数权重。为实现高效训练,我们推导了不仅针对MPC模块,还针对基于优化的门框穿越检测模块的解析策略梯度。硬件实验表明,系统可实现峰值加速度达30 m/s²的敏捷穿门,并在体感速率扰动超过1146 deg/s的情况下于0.85秒内完成恢复。
原文摘要 · Abstract (English)
Traversing narrow gates presents a significant challenge and has become a standard benchmark for evaluating agile and precise quadrotor flight. Traditional modularized autonomous flight stacks require extensive design and parameter tuning, while end-to-end reinforcement learning (RL) methods often suffer from low sample efficiency, limited interpretability, and degraded disturbance rejection under unseen perturbations. In this work, we present a novel hybrid framework that adaptively fine-tunes model predictive control (MPC) parameters online using outputs from a neural network (NN) trained offline. The NN jointly predicts a reference pose and cost function weights, conditioned on the coordinates of the gate corners and the current drone state. To achieve efficient training, we derive analytical policy gradients not only for the MPC module but also for an optimization-based gate traversal detection module. Hardware experiments demonstrate agile and accurate gate traversal with peak accelerations of $30\ \mathrm{m/s^2}$, as well as recovery within $0.85\ \mathrm{s}$ following body-rate disturbances exceeding $1146\ \mathrm{deg/s}$.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。