利用机器人系统的对称性,加速轨迹跟踪控制器的强化学习训练。
Leveraging Symmetry to Accelerate Learning of Trajectory Tracking Controllers for Free-Flying Robotic Systems
- 基于李群对称性构建低维商空间马尔可夫决策过程
- 训练速度提升3倍,收敛后跟踪误差降低40%
- 适合高维、复杂动力学的飞行机器人控制设计
轨迹跟踪控制器使机器人系统能精确跟随规划的参考轨迹。强化学习(RL)在具有复杂动力学且在线计算资源有限的系统中展现潜力。然而,RL样本效率低和奖励函数设计困难导致训练缓慢且不稳定,尤其在高维系统中。本文利用具有浮游基底的机器人系统的内在李群对称性,缓解这些挑战。我们将一般跟踪问题建模为包含物理状态与参考状态演化的马尔可夫决策过程(MDP)。随后证明,底层动力学和运行代价中的对称性导致一个MDP同态,该映射允许在低维“商”MDP上训练的策略被提升为原系统的最优跟踪控制器。我们使用近端策略优化(PPO)方法,在三个系统上比较了该对称性感知方法与无结构基线:粒子(受力质点)、Astrobee(全驱动空间机器人)和四旋翼(欠驱动系统)。结果表明,对称性感知方法不仅显著加速训练,还降低了收敛时的跟踪误差。
原文摘要 · Abstract (English)
Tracking controllers enable robotic systems to accurately follow planned reference trajectories. In particular, reinforcement learning (RL) has shown promise in the synthesis of controllers for systems with complex dynamics and modest online compute budgets. However, the poor sample efficiency of RL and the challenges of reward design make training slow and sometimes unstable, especially for high-dimensional systems. In this work, we leverage the inherent Lie group symmetries of robotic systems with a floating base to mitigate these challenges when learning tracking controllers. We model a general tracking problem as a Markov decision process (MDP) that captures the evolution of both the physical and reference states. Next, we prove that symmetry in the underlying dynamics and running costs leads to an MDP homomorphism, a mapping that allows a policy trained on a lower-dimensional "quotient" MDP to be lifted to an optimal tracking controller for the original system. We compare this symmetry-informed approach to an unstructured baseline, using Proximal Policy Optimization (PPO) to learn tracking controllers for three systems: the Particle (a forced point mass), the Astrobee (a fullyactuated space robot), and the Quadrotor (an underactuated system). Results show that a symmetry-aware approach both accelerates training and reduces tracking error at convergence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。