用Transformer提升强化学习效率,优化电动垂直起降无人机的省电起飞轨迹。
Transformer-Guided Deep Reinforcement Learning for Optimal Takeoff Trajectory Design of an eVTOL Drone
- 引入Transformer动态分析状态空间,降低强化学习训练难度。
- 仅需原方法25%的训练步数(457万对比1979万),能耗逼近最优解。
- 适合关注智能飞行控制与低空交通的科研与工程人员。
电动垂直起降(eVTOL)飞行器的快速发展为缓解城市交通拥堵提供了可能,但其起飞阶段过高的能耗仍是应用瓶颈。为实现最小能耗的最优起飞轨迹设计,传统最优控制方法受限于高维复杂性,而深度强化学习(DRL)又面临训练困难。本文提出基于Transformer引导的DRL方法,通过在每一步探索真实状态空间,显著降低训练难度。该方法在eVTOL无人机起飞轨迹优化中验证,通过调节功率与机翼倾角控制变量,在满足最小垂直位移和水平速度条件下实现节能目标。结果表明,所提方法仅需4.57×10⁶次训练时间步,仅为原始DRL代理所需19.79×10⁶步的25%;同时在能耗准确性上达到97.2%,优于原始DRL的96.1%。因此,该方法在训练效率与最优设计验证方面均优于传统DRL。
原文摘要 · Abstract (English)
The rapid advancement of electric vertical takeoff and landing (eVTOL) aircraft offers a promising opportunity to alleviate urban traffic congestion but is still limited by excessive power demands, especially during the takeoff phase. Thus, developing optimal takeoff trajectories for minimum energy consumption becomes essential for broader eVTOL aircraft applications. Conventional optimal control methods (such as dynamic programming and linear quadratic regulator) provide highly efficient and well-established solutions but are prohibited by problem dimensionality and complexity. Deep reinforcement learning (DRL) emerges as a special type of artificial intelligence tackling complex, nonlinear systems; however, the training difficulty is a key bottleneck that hinders DRL applications. To address these challenges, we propose the transformer-guided DRL to alleviate the training difficulty by exploring a realistic state space at each time step using a transformer. The proposed transformer-guided DRL was demonstrated on an optimal takeoff trajectory design of an eVTOL drone for minimal energy consumption while meeting takeoff conditions (i.e., minimum vertical displacement and minimum horizontal velocity) by varying control variables (i.e., power and wing angle to the vertical). Results presented that the transformer-guided DRL agent learned to take off with $4.57\times10^6$ time steps, representing $25\%$ of the $19.79\times10^6$ time steps needed by a vanilla DRL agent. In addition, the transformer-guided DRL achieved $97.2\%$ accuracy on the optimal energy consumption compared against the simulation-based optimal reference, while the vanilla DRL achieved $96.1\%$ accuracy. Therefore, the proposed transformer-guided DRL outperformed vanilla DRL in terms of both training efficiency and optimal design verification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。