arXiv:2512.13359cs.ROcs.LG2025-12被引 4

用加速强化学习实现水下机器人六自由度快速精准控制

Fast Policy Learning for 6-DOF Position Control of Underwater Vehicles

  • 基于JAX和MJX的并行仿真与学习联合编译,训练仅需两分钟
  • 真实水下实验中实现零样本迁移的稳定轨迹跟踪与抗干扰能力
  • 适合需要快速部署高精度水下控制策略的研究者与工程师

自主水下航行器(AUV)在复杂动态海洋环境中高效运行需可靠的六自由度(6-DOF)位置控制。传统控制器在理想条件下有效,但面对未建模动态或环境扰动时性能下降。强化学习(RL)提供了有力替代方案,但训练通常缓慢且存在仿真到现实的迁移难题。本文提出基于JAX和MuJoCo-XLA(MJX)的GPU加速强化学习训练流水线。通过联合即时编译大规模并行物理仿真与学习更新,实现训练时间低于两分钟。通过系统评估多种强化学习算法,展示了在真实水下实验中稳健的6-DOF轨迹跟踪与有效的扰动抑制能力,且策略可实现零样本从仿真到现实的迁移。

原文摘要 · Abstract (English)

Autonomous Underwater Vehicles (AUVs) require reliable six-degree-of-freedom (6-DOF) position control to operate effectively in complex and dynamic marine environments. Traditional controllers are effective under nominal conditions but exhibit degraded performance when faced with unmodeled dynamics or environmental disturbances. Reinforcement learning (RL) provides a powerful alternative but training is typically slow and sim-to-real transfer remains challenging. This work introduces a GPU accelerated RL training pipeline built in JAX and MuJoCo-XLA (MJX). By jointly JIT-compiling large-scale parallel physics simulation and learning updates, we achieve training times of under two minutes. Through systematic evaluation of multiple RL algorithms, we show robust 6-DOF trajectory tracking and effective disturbance rejection in real underwater experiments, with policies transferred zero-shot from simulation.

强化学习水下控制仿真迁移实时控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。