arXiv:2502.12355cs.ROcs.LG2025-02中稿 · 2025 IEEE Internat…

用深度强化学习实现软驱动微型无人机零样本悬停飞行

Hovering Flight of Soft-Actuated Insect-Scale Micro Aerial Vehicles using Deep Reinforcement Learning

  • 结合改进行为克隆与强化学习,解决微尺度飞行延迟和不确定性问题
  • 实测最长悬停50秒,横向误差1.34厘米,垂直误差0.05厘米
  • 首次在两种720-850mg软驱动微型飞行器上实现端到端深度强化学习飞行

软驱动的昆虫尺度微型飞行器(IMAVs)在设计鲁棒且计算高效的控制器方面面临独特挑战。在毫米尺度下,系统动态响应极快(约毫秒级),叠加系统延迟、模型不确定性及外部扰动,显著影响飞行性能。本文设计了一种深度强化学习(RL)控制器,以应对延迟与不确定性。为初始化神经网络控制器,提出一种改进的行为克隆(BC)方法,结合状态-动作重匹配处理延迟,并采用领域随机化的专家示范以应对不确定性。随后应用近端策略优化(PPO)进行强化学习微调,提升性能并平滑控制指令。仿真结果表明,改进后的BC相比基线提升平均奖励;采用PPO的强化学习显著改善飞行质量并减少指令波动。该控制器部署于两台重量分别为720 mg和850 mg的昆虫尺度飞行机器人上,均成功实现多次零样本悬停飞行,最长持续50秒,横向均方根误差为1.34厘米,垂直方向为0.05厘米,标志着首个基于深度强化学习的端到端飞行控制在软驱动IMAV上的实现。

原文摘要 · Abstract (English)

Soft-actuated insect-scale micro aerial vehicles (IMAVs) pose unique challenges for designing robust and computationally efficient controllers. At the millimeter scale, fast robot dynamics ($\sim$ms), together with system delay, model uncertainty, and external disturbances significantly affect flight performances. Here, we design a deep reinforcement learning (RL) controller that addresses system delay and uncertainties. To initialize this neural network (NN) controller, we propose a modified behavior cloning (BC) approach with state-action re-matching to account for delay and domain-randomized expert demonstration to tackle uncertainty. Then we apply proximal policy optimization (PPO) to fine-tune the policy during RL, enhancing performance and smoothing commands. In simulations, our modified BC substantially increases the mean reward compared to baseline BC; and RL with PPO improves flight quality and reduces command fluctuations. We deploy this controller on two different insect-scale aerial robots that weigh 720 mg and 850 mg, respectively. The robots demonstrate multiple successful zero-shot hovering flights, with the longest lasting 50 seconds and root-mean-square errors of 1.34 cm in lateral direction and 0.05 cm in altitude, marking the first end-to-end deep RL-based flight on soft-driven IMAVs.

强化学习微型飞行器软体驱动端到端控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。