arXiv:2502.20500cs.ROmath.OC2025-02被引 5

通过对称性建模提升无人机强化学习的采样效率与泛化能力。

Equivariant Reinforcement Learning Frameworks for Quadrotor Low-Level Control

  • 利用无人机动力学的旋转与反射对称性,构建等变网络模型。
  • 在相同训练数据下,飞行性能提升,学习效率显著优于传统方法。
  • 适合需要高效安全训练的低层无人机控制场景。

提升采样效率和泛化能力对于不稳定的四旋翼无人飞行器(UAV)数据驱动控制至关重要。尽管已有多种强化学习(RL)方法应用于自主四旋翼飞行,但通常需要大量训练数据,实际应用中面临多重挑战与安全风险。为此,我们提出数据高效、等变的统一与模块化强化学习框架,用于四旋翼低层控制。具体而言,通过识别四旋翼动力学中的旋转与反射对称性,并将这些对称性编码进等变网络模型,消除状态-动作空间中的冗余学习。该方法使某一构型下学习到的最优控制动作能自动推广至其他构型,从而提升数据效率。实验结果表明,我们的等变方法在学习效率与飞行性能上显著优于非等变基线方法。

原文摘要 · Abstract (English)

Improving sampling efficiency and generalization capability is critical for the successful data-driven control of quadrotor unmanned aerial vehicles (UAVs) that are inherently unstable. While various reinforcement learning (RL) approaches have been applied to autonomous quadrotor flight, they often require extensive training data, posing multiple challenges and safety risks in practice. To address these issues, we propose data-efficient, equivariant monolithic and modular RL frameworks for quadrotor low-level control. Specifically, by identifying the rotational and reflectional symmetries in quadrotor dynamics and encoding these symmetries into equivariant network models, we remove redundancies of learning in the state-action space. This approach enables the optimal control action learned in one configuration to automatically generalize into other configurations via symmetry, thereby enhancing data efficiency. Experimental results demonstrate that our equivariant approaches significantly outperform their non-equivariant counterparts in terms of learning efficiency and flight performance.

强化学习无人机控制等变网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。