一种可即插即用的实时强化学习算法,提升仿生双翼飞行器在复杂环境下的稳定控制能力。
A Plug-and-Play Fully On-the-Job Real-Time Reinforcement Learning Algorithm for a Direct-Drive Tandem-Wing Experiment Platforms Under Multiple Random Operating Conditions
- 结合物理启发规则与扰动模块,设计轻量网络实现全流程在线学习。
- 500步内完成稳定训练,追踪精度比主流算法提升14至66倍。
- 适合对实时性与鲁棒性要求高的无人机控制场景,尤其在突变环境下表现优异。
此类仿生系统中双翼产生的非线性不稳定性气动干扰给运动控制带来巨大挑战,尤其是在多种随机运行条件下。为此,本文提出CRL2E算法——一种可即插即用、全在线实时强化学习算法,融合了新颖的物理启发式规则驱动策略组合策略与扰动模块,并采用面向实时控制优化的轻量化网络结构。为验证性能与模块设计合理性,在六种严苛运行条件下对比了七种不同算法。结果表明,CRL2E算法在前500步内即可实现安全稳定的训练,相较于Soft Actor-Critic、Proximal Policy Optimization和Twin Delayed Deep Deterministic Policy Gradient算法,追踪精度提升14至66倍;在各类随机运行条件下,相比CRL算法,追踪精度提升8.3%至60.4%。其收敛速度比仅含组合扰动的CRL算法快36.11%至57.64%,比同时引入组合与时间交错扰动的CRL算法快43.52%至65.85%,尤其在标准CRL难以收敛的条件下优势显著。硬件测试显示,优化后的轻量化网络结构在负载重量与平均推理时间方面表现优异,满足实时控制需求。
原文摘要 · Abstract (English)
The nonlinear and unstable aerodynamic interference generated by the tandem wings of such biomimetic systems poses substantial challenges for motion control, especially under multiple random operating conditions. To address these challenges, the Concerto Reinforcement Learning Extension (CRL2E) algorithm has been developed. This plug-and-play, fully on-the-job, real-time reinforcement learning algorithm incorporates a novel Physics-Inspired Rule-Based Policy Composer Strategy with a Perturbation Module alongside a lightweight network optimized for real-time control. To validate the performance and the rationality of the module design, experiments were conducted under six challenging operating conditions, comparing seven different algorithms. The results demonstrate that the CRL2E algorithm achieves safe and stable training within the first 500 steps, improving tracking accuracy by 14 to 66 times compared to the Soft Actor-Critic, Proximal Policy Optimization, and Twin Delayed Deep Deterministic Policy Gradient algorithms. Additionally, CRL2E significantly enhances performance under various random operating conditions, with improvements in tracking accuracy ranging from 8.3% to 60.4% compared to the Concerto Reinforcement Learning (CRL) algorithm. The convergence speed of CRL2E is 36.11% to 57.64% faster than the CRL algorithm with only the Composer Perturbation and 43.52% to 65.85% faster than the CRL algorithm when both the Composer Perturbation and Time-Interleaved Capability Perturbation are introduced, especially in conditions where the standard CRL struggles to converge. Hardware tests indicate that the optimized lightweight network structure excels in weight loading and average inference time, meeting real-time control requirements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。