arXiv:2607.01528cs.LGcs.SY2026-07

用学习到的风速估计提升小型无人机在强风中的飞行控制精度。

Wind-Aware Reinforcement Learning Control of a Small Quadrotor Using Learned Onboard Wind Estimation in Simulated Atmospheric Turbulence

论文配图:Wind-Aware Reinforcement Learning Control of a Small Quadrotor Using Learned Onboard Wind Estimation in Simulated Atmospheric Turbulence
图 1 · 摘自论文原文
  • 通过神经网络从飞行数据中实时估算风速,精度达0.40 m/s
  • 强化学习控制器使水平轨迹误差降低48%,强风下仍稳定有效
  • 适合研究无人机自主飞行与复杂环境控制的学者和工程师

小型多旋翼飞行器越来越多地被用于大气边界层作业,那里的湍流风速与飞行器空速相当,会严重破坏轨迹跟踪并超越传统反馈控制能力。本文提出一种两阶段学习流程:首先利用机载运动学与动力学数据,通过注意力增强的门控循环网络估计局部风速;该模型在数千次模拟飞行中训练,覆盖冯·卡门湍流、幂律剪切和偏转,对未见过的风场实现每飞行平均均方根误差0.40 m/s、方向误差3.2度,接近由未解析湍流决定的理论下限;在垂直爬升中,其技能得分达0.861(相对于恒定风参考)。随后,一个采用近端策略优化的强化学习飞行控制器接收冻结的风估计输入,在平均风速4–12 m/s范围内,相比无风感知的PD基线,水平轨迹误差减少48%,所有评估回合均获胜。三重消融分析表明,性能提升来自运动学成分(无需风信息)与风感知成分,后者随风速增加而上升,在强风中约占总收益的一半,符合气动阻力的平方量级增长规律。当遭遇13–15 m/s分布外风时,控制器仍能优雅退化,而基线则完全失效。

原文摘要 · Abstract (English)

Small multirotor aircraft are increasingly tasked with operations in the atmospheric boundary layer, where turbulent winds comparable to the vehicle's airspeed degrade trajectory tracking and can defeat conventional feedback control. This work illustrates a two-stage learning pipeline that first estimates the local wind from onboard kinematics and dynamics and then exploits that estimate inside a reinforcement learning (RL) flight controller. The wind estimator, an attention-augmented gated recurrent network trained on thousands of simulated flights through von Karman turbulence with power-law shear and veer, recovers the horizontal wind vector with a per-flight root-mean-square error of 0.40 m/s and a direction error of 3.2 degrees on unseen wind regimes, an accuracy near the floor imposed by unresolved turbulence, and generalizes to vertical ascent profiles with a skill score of 0.861 over a constant-wind reference. A proximal policy optimization controller receiving the frozen estimator's output reduces horizontal trajectory tracking error by 48% relative to a wind-blind proportional-derivative baseline across mean winds of 4 m/s to 12 m/s, winning on 100% of evaluation episodes. A three-way ablation decomposes this improvement into a kinematic component, available without wind information, and a wind-perception component; the perception share rises with wind speed, from small in light winds toward roughly half the total benefit in strong winds, consistent with the quadratic scaling of aerodynamic drag. The controller degrades gracefully on out-of-distribution winds of 13 m/s to 15 m/s, where the baseline fails catastrophically.

无人机控制强化学习风估计鲁棒飞行

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。