用数字孪生+强化学习,让水下机器人避障更智能。
Digital Twin Supervised Reinforcement Learning Framework for Autonomous Underwater Navigation
- 基于PPO算法融合目标信息与虚拟栅格,实现水下自主导航。
- 在复杂障碍环境中碰撞率显著低于传统规划方法。
- 仿真训练结果可直接迁移到真实水下机器人上。
水下自主导航因缺乏GPS、能见度低及存在潜藏障碍物而面临重大挑战。本文以广泛用于科学实验的BlueROV2为研究对象,提出一种基于近端策略优化(PPO)的深度强化学习方法,其观测空间融合了目标导向信息、虚拟占用栅格及操作区域边界的射线扫描数据。所学策略与常用的动态窗口法(DWA)——一种稳健的避障基准方法——进行对比。评估在逼真的仿真环境中进行,并通过物理BlueROV2在3D数字孪生测试场景下的验证得以补充,有效降低真实实验风险。结果显示,该PPO策略在高度杂乱环境中持续优于DWA,尤其体现在局部适应能力更强且碰撞更少。最终实验验证了从仿真到现实世界的策略可迁移性,证实深度强化学习在水下机器人自主导航中的有效性。
原文摘要 · Abstract (English)
Autonomous navigation in underwater environments remains a major challenge due to the absence of GPS, degraded visibility, and the presence of submerged obstacles. This article investigates these issues through the case of the BlueROV2, an open platform widely used for scientific experimentation. We propose a deep reinforcement learning approach based on the Proximal Policy Optimization (PPO) algorithm, using an observation space that combines target-oriented navigation information, a virtual occupancy grid, and ray-casting along the boundaries of the operational area. The learned policy is compared against a reference deterministic kinematic planner, the Dynamic Window Approach (DWA), commonly employed as a robust baseline for obstacle avoidance. The evaluation is conducted in a realistic simulation environment and complemented by validation on a physical BlueROV2 supervised by a 3D digital twin of the test site, helping to reduce risks associated with real-world experimentation. The results show that the PPO policy consistently outperforms DWA in highly cluttered environments, notably thanks to better local adaptation and reduced collisions. Finally, the experiments demonstrate the transferability of the learned behavior from simulation to the real world, confirming the relevance of deep RL for autonomous navigation in underwater robotics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。