arXiv:2605.16015cs.ROcs.LG2026-05

用强化学习让无人机实时感知并自适应抗干扰,飞得更稳更准。

Adaptive Outer-Loop Control of Quadrotors via Reinforcement Learning

论文配图:Adaptive Outer-Loop Control of Quadrotors via Reinforcement Learning
图 1 · 摘自论文原文
  • 用残差动力预测器在线估算扰动,替代依赖真实数据的旧方法
  • 实测在质量变化、偏载和动态吊挂下仍能精准追踪轨迹
  • 仅需几秒飞行数据即可完成仿真到现实的校准,适合小型无人机部署

针对四旋翼飞行控制中的深度强化学习(DRL)因依赖领域随机化(DR)导致策略过于保守、难以应对动态扰动的问题,本文提出一种自适应外环控制架构。首先训练最优外环策略,再以残差动力预测器(RDP)取代对真实扰动数据的依赖,RDP仅通过历史状态与控制动作在线估计飞行中外部受力与力矩。为实现无缝硬件迁移,引入数据高效线性校准桥与在线推力修正机制,仅需数秒飞行数据即可对齐仿真隐空间与真实环境。在Crazyflie微型四旋翼上的真实世界验证表明,所提自适应控制器显著优于基线方法,在质量变化、非对称负载及动态吊挂等严重不确定性条件下仍能保持精确轨迹跟踪。

原文摘要 · Abstract (English)

Deep Reinforcement Learning (DRL) for quadrotor flight control typically relies on Domain Randomization (DR) for sim-to-real transfer, resulting in overly conservative policies that struggle with dynamic disturbances. To overcome this, we propose a novel adaptive control architecture that actively perceives and reacts to instantaneous perturbations. First, we train an optimal outer-loop policy, then replace its reliance on ground-truth disturbance data with a Residual Dynamics Predictor (RDP). The RDP estimates the external forces and moments acting on the aircraft in flight online using only the history of states and control actions. For seamless hardware transfer, we introduce a data-efficient linear calibration bridge and an online thrust correction mechanism that align the simulated latent space with reality using mere seconds of flight data. Real-world validations on a Crazyflie micro-quadrotor demonstrate that our adaptive controller significantly outperforms baselines, maintaining precise trajectory tracking under severe uncertainties including mass variations, asymmetric payloads, and dynamic slung loads

强化学习无人机控制自适应控制端到端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。