利用旋转对称性提升无人机特技飞行的强化学习效率与泛化能力
Multi-Task Reinforcement Learning of Drone Aerobatics by Exploiting Geometric Symmetries
- 基于SO(2)旋转对称性设计等变策略网络,显式建模动力学特性
- 在多种特技任务中达到98.85%成功率,显著优于基线方法
- 适合需要统一控制多种复杂机动的自主无人机系统研发
自主微型飞行器(MAVs)的飞行控制正从近平衡点的平稳飞行转向更激进的特技动作,如翻滚、横滚和动力环飞。尽管强化学习(RL)在这些任务中展现出巨大潜力,但传统方法普遍存在数据效率低、泛化能力差的问题,尤其在多任务场景下更为突出。本文提出一种端到端的多任务强化学习框架GEAR(Geometric Equivariant Aerobatics Reinforcement),充分挖掘MAV动力学中的固有SO(2)旋转对称性,并将其显式嵌入策略网络架构。通过引入等变演员网络、基于FiLM的任务调制机制以及多头评论家,GEAR实现了多样特技动作的高效且灵活学习,构建了一个数据高效、鲁棒性强且统一的特技控制框架。GEAR在多种特技任务中实现98.85%的成功率,显著超越基线方法。真实实验表明,GEAR能稳定执行多种动作,并可组合基础运动原语完成复杂特技。
原文摘要 · Abstract (English)
Flight control for autonomous micro aerial vehicles (MAVs) is evolving from steady flight near equilibrium points toward more aggressive aerobatic maneuvers, such as flips, rolls, and Power Loop. Although reinforcement learning (RL) has shown great potential in these tasks, conventional RL methods often suffer from low data efficiency and limited generalization. This challenge becomes more pronounced in multi-task scenarios where a single policy is required to master multiple maneuvers. In this paper, we propose a novel end-to-end multi-task reinforcement learning framework, called GEAR (Geometric Equivariant Aerobatics Reinforcement), which fully exploits the inherent SO(2) rotational symmetry in MAV dynamics and explicitly incorporates this property into the policy network architecture. By integrating an equivariant actor network, FiLM-based task modulation, and a multi-head critic, GEAR achieves both efficiency and flexibility in learning diverse aerobatic maneuvers, enabling a data-efficient, robust, and unified framework for aerobatic control. GEAR attains a 98.85\% success rate across various aerobatic tasks, significantly outperforming baseline methods. In real-world experiments, GEAR demonstrates stable execution of multiple maneuvers and the capability to combine basic motion primitives to complete complex aerobatics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。