arXiv:2603.12020cs.ROcs.AI2026-03

用高保真仿真加速水下机器人自主对接训练,实测成功率超90%。

Sim-to-reality adaptation for Deep Reinforcement Learning applied to an underwater docking application

  • 基于数字孪生构建多进程强化学习框架,加速训练并模拟真实水下环境。
  • 在仿真中实现90%以上对接成功率,物理测试验证有效且出现自适应对齐行为。
  • 适合水下机器人控制、强化学习落地应用的研究者参考。

深度强化学习(DRL)为自主水下对接提供了一种鲁棒替代方案,尤其适用于应对不可预测的环境条件。然而,弥合“仿真到现实”的差距以及高训练延迟仍是实际部署的主要瓶颈。本文针对吉罗纳自主水下航行器(Girona AUV),提出系统性方法,利用高保真数字孪生环境实现自主对接。将Stonefish仿真器改造为多进程强化学习框架,显著加速学习过程,同时融入真实的AUV动力学、碰撞模型和传感器噪声。采用近端策略优化(PPO)算法,在无头环境中训练6-自由度控制策略,并通过随机起始位置提升泛化能力。奖励函数综合考虑距离、姿态、动作平滑性及自适应碰撞惩罚,以实现软对接。实验结果表明,该智能体在仿真中成功率达90%以上。此外,物理测试水池验证了仿真到现实迁移的有效性,DRL控制器展现出如基于俯仰的制动和偏航振荡等自适应对齐的涌现行为。

原文摘要 · Abstract (English)

Deep Reinforcement Learning (DRL) offers a robust alternative to traditional control methods for autonomous underwater docking, particularly in adapting to unpredictable environmental conditions. However, bridging the "sim-to-real" gap and managing high training latencies remain significant bottlenecks for practical deployment. This paper presents a systematic approach for autonomous docking using the Girona Autonomous Underwater Vehicle (AUV) by leveraging a high-fidelity digital twin environment. We adapted the Stonefish simulator into a multiprocessing RL framework to significantly accelerate the learning process while incorporating realistic AUV dynamics, collision models, and sensor noise. Using the Proximal Policy Optimization (PPO) algorithm, we developed a 6-DoF control policy trained in a headless environment with randomized starting positions to ensure generalized performance. Our reward structure accounts for distance, orientation, action smoothness, and adaptive collision penalties to facilitate soft docking. Experimental results demonstrate that the agent achieved a success rate of over 90% in simulation. Furthermore, successful validation in a physical test tank confirmed the efficacy of the sim-to-reality adaptation, with the DRL controller exhibiting emergent behaviors such as pitch-based braking and yaw oscillations to assist in mechanical alignment.

强化学习水下机器人仿真迁移数字孪生

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。