无需微调,直接将仿真训练的控制策略迁移到真实小车摆杆系统。
Zero-shot Transfer of Reinforcement Learning Control Policies for the Swing-Up and Stabilization of a Cart-Pole System

- 独立训练起摆与稳定策略,通过逻辑切换实现无缝交接。
- 在所有实验中,起摆策略均成功将摆杆送入稳定器吸引域。
- 结合滤波与随机化技术,适合高鲁棒性控制场景应用。
强化学习(RL)是现代化控制器设计的强大且便捷工具。本文研究了基于强化学习的控制策略从仿真到硬件的零样本迁移,应用于小车-摆杆系统的起摆与稳定任务。两个策略独立训练,通过Simulink中的切换逻辑实现交接。采用一阶动作平滑滤波器,防止高频振荡导致硬件损坏。结合感知敏感性的领域随机化(DR)与简单的线性课程学习(CL)调度,所有实验中起摆策略均能注入足够能量,使系统进入稳定器的吸引域。稳定策略可有效抑制测试范围内的扰动,起摆策略在更大扰动后仍可重新激活,并将摆杆恢复至倒立位置。
原文摘要 · Abstract (English)
Reinforcement learning (RL) is a powerful and convenient tool to modernize controller design. In this work, we study the zero-shot transfer of RL-based control policies from simulation to hardware for cart-pole swing-up and stabilization. The two policies are trained independently, and the handoff is implemented in Simulink via switching logic. We apply a first-order action smoothing filter to prevent hardware damage from high-frequency oscillatory actuation. Pairing this bandwidth-aware filtering with sensitivity-guided domain randomization (DR) and a simple linear curriculum learning (CL) schedule, we obtain a swing-up policy that in all of our experiments injects sufficient energy for handoff into the stabilizer's region of attraction. The stabilization policy rejects disturbances within the tested range, and the swing-up policy can re-engage after larger perturbations and restores the pendulum to the inverted position.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。