用强化学习实现水下机器人6自由度快速控制,无需调参直接落地
Learning to Swim: Reinforcement Learning for 6-DOF Control of Thruster-driven Autonomous Underwater Vehicles
- 基于强化学习直接映射6自由度指令到推进器输出
- 训练仅需数分钟,仿真到实物零样本迁移效果媲美人工调优
- 模拟器支持领域随机化,对物理参数微小变化具有鲁棒性
水下机器人控制因复杂的非线性水动力作用而困难,尤其对小型机器人,负载与环境变化会显著改变动力学特性。传统方法依赖被动稳定(上下配重)和手工调优的PID控制器,但响应迟缓且需频繁重调。本文提出一种可快速训练(数分钟内完成)的强化学习方法,实现全6自由度推进式水下机器人的直接控制,输入为6自由度指令,输出为推进器命令。我们构建了一个高度并行化的水下动力学仿真器,并在真实水下机器人上实现零样本仿真到现实迁移(无任何调参),性能达到人工调优PID水平。此外,通过模拟器中的领域随机化,训练出的策略对车辆物理参数的小幅变化具有鲁棒性。
原文摘要 · Abstract (English)
Controlling AUVs can be challenging because of the effect of complex non-linear hydrodynamic forces acting on the robot, which are significant in water and cannot be ignored. The problem is exacerbated for small AUVs for which the dynamics can change significantly with payload changes and deployments under different hydrodynamic conditions. The common approach to AUV control is a combination of passive stabilization with added buoyancy on top and weights on the bottom, and a PID controller tuned for simple and smooth motion primitives. However, the approach comes at the cost of sluggish controls and often the need to re-tune controllers with configuration changes. In this paper, we propose a fast (trainable in minutes), reinforcement learning-based approach for full 6 degree of freedom (DOF) control of a thruster-driven AUVs, taking 6-DOF command-conditioned inputs direct to thruster outputs. We present a new, highly parallelized simulator for underwater vehicle dynamics. We demonstrate this approach through zero-shot sim-to-real (with no tuning) transfer onto a real AUV that produces comparable results to hand-tuned PID controllers. Furthermore, we show that domain randomization on the simulator produces policies that are robust to small variations in vehicle's physical parameters.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。