通过奖励设计与终止条件,实现无人机控制性能的可调性。
A Heuristic Approach for Performance Tuning in RL-based Quadrotor Control via Reward Design and Termination Conditions

- 设计双带宽指数奖励结构,实现基准临界阻尼响应。
- 600万步内样本高效训练,稳态误差约2%。
- 调整参数即可实现快(特技)或慢(巡检)的响应性能。
基于强化学习的四轴飞行器控制在复杂环境快速导航和无人机竞速等任务中表现出色,但在基础设施巡检等场景中,精准、可控且可调的运动性能至关重要。本文提出一种新颖的启发式方法,通过奖励设计与终止条件调控性能。设计包含双带宽指数项的奖励结构,在PPO算法配合时段截断条件下,仅用600万时间步即可实现样本高效的基准临界阻尼响应,稳态误差约2%。为调节性能,提出直观的启发式规则,通过调整奖励权重与指数系数,可实现更快(特技类)或更慢(巡检类)的收敛时间,同时保持基准临界阻尼特性与近似2%的稳态误差。在100次随机初始条件测试中,三种策略均展现出准确的位置与偏航跟踪能力,验证了该方法的有效性。
原文摘要 · Abstract (English)
Reinforcement learning (RL)-based quadrotor control policies have achieved impressive performance in tasks such as fast navigation in cluttered environments and drone racing, where the focus is on speed and agility. However, in several applications, such as infrastructure inspection, it is critical to achieve precise, controlled maneuvers with tunable performance. In this article, we present a novel heuristic approach to achieve tunable performance in RL-based Quadrotor control through reward design and termination conditions. We present a novel reward structure containing dual bandwidth exponentials that achieves a baseline critically damped response in setpoint tracking, with low steady-state errors. When trained with a Proximal Policy Optimization (PPO) algorithm, in conjunction with episode truncation conditions, the desired performance is achieved in 6 million time steps in a sample-efficient manner. In order to tune the performance about the baseline behavior, we present intuitive heuristic rules to adjust the reward weights and exponential coefficients to achieve faster (acrobatic-like) and slower (inspection-like) settling time performance, while retaining the baseline critically damped response and approximately 2\% steady-state error. We evaluate the three RL policies (baseline, acrobatic, and inspection) across 100 trials and show accurate and tunable performance in position and yaw tracking from random initial conditions, thereby demonstrating the effectiveness of the proposed heuristic approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。