用强化学习+非线性控制,让小无人机在强风中精准飞行。
GustPilot: A Hierarchical DRL-INDI Framework for Wind-Resilient Quadrotor Navigation
- 分层设计:强化学习规划速度,非线性控制器快速抗风扰动。
- 真实飞行测试中成功率94.7%,远超基线55.0%;风速达3.5m/s仍能稳定飞行。
- 无需重新训练即可适应多门多风扇复杂环境,适合实际部署的无人机系统。
风扰动仍是轻量级四旋翼自主导航的主要障碍,快速变化的气流会破坏规划与跟踪的稳定性。本文提出GustPilot,一种分层风阻抗导航框架:深度强化学习(DRL)策略生成惯性坐标系下的速度参考以穿越门框;同时,基于几何增量非线性动态逆(INDI)的低层控制器利用机载传感器测量值,对线加速度和角加速度率进行增量反馈,实现快速残差扰动抑制。通过双层策略保障鲁棒性:训练阶段采用风扇喷流域随机化进行风感知规划,运行时由INDI控制器实现即时扰动抵消。我们在50g四旋翼平台上实测了四种场景(无风至动态风场、移动门框、移动扰动源),对比了DRL-PID基线。尽管仅在单门单风扇环境中训练,该策略可泛化至最多六门四风扇的复杂环境而无需重训。80次实验中,DRL-INDI平均整体成功率达94.7%,显著优于基线55.0%;跟踪均方根误差降低最高达50%,在风速高达3.5 m/s时仍能维持1.34 m/s的飞行速度。结果表明,结合DRL速度规划与结构化INDI抗扰机制,是一种实用且可泛化的风阻抗自主飞行方案。
原文摘要 · Abstract (English)
Wind disturbances remain a key barrier to reliable autonomous navigation for lightweight quadrotors, where the rapidly varying airflow can destabilize both planning and tracking. This paper introduces GustPilot, a hierarchical wind-resilient navigation stack in which a deep reinforcement learning (DRL) policy generates inertial-frame velocity reference for gate traversal. At the same time, a geometric Incremental Nonlinear Dynamic Inversion (INDI) controller provides low-level tracking with fast residual disturbance rejection. The INDI layer achieves this by providing incremental feedback on both specific linear acceleration and angular acceleration rate, using onboard sensor measurements to reject wind disturbances rapidly. Robustness is obtained through a two-level strategy, wind-aware planning learned via fan-jet domain randomization during training, and rapid execution-time disturbance rejection by the INDI tracking controller. We evaluate GustPilot in real flights on a 50g quad-copter platform against a DRL-PID baseline across four scenarios ranging from no-wind to fully dynamic conditions with a moving gate and a moving disturbance source. Despite being trained only in a minimal single-gate and single-fan setup, the policy generalizes to significantly more complex environments (up to six gates and four fans) without retraining. Across 80 experiments, DRL-INDI achieves a 94.7% versus 55.0% for DRL-PID as average Overall Success Rate (OSR), reduces tracking RMSE up to 50%, and sustains speeds up to 1.34 m/s under wind disturbances up to 3.5 m/s. These results demonstrate that combining DRL-based velocity planning with structured INDI disturbance rejection provides a practical and generalizable approach to wind-resilient autonomous flight navigation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。