用强化学习优化船舶在复杂水道中的智能导航,支持多起点终点和不同网格精度。
Goal-Conditioned Reinforcement Learning for Data-Driven Maritime Navigation
- 基于强化学习与交通图数据,实现多路径智能导航决策。
- 引入正向奖励与惩罚机制,提升航行效率与路线多样性。
- 适用于复杂水域的船舶自主导航,适合航运智能化研究者。
在狭窄且动态变化的水道中规划船舶航线极具挑战,因环境条件不断变化及运营约束限制。现有研究通常无法跨多个起点-终点对泛化,且未充分利用大规模数据驱动的交通图。本文提出一种针对大型海事数据的强化学习方法,可学习在多个起点-终点间规划路径,并适应不同六边形网格分辨率。智能体在连续观测下于多离散动作空间中选择方向与速度。奖励函数综合考虑燃油效率、航行时间、风阻及路线多样性,基于由自动识别系统(AIS)构建的交通图与ERA5风场数据。方法在世界最大河口之一——圣劳伦斯湾进行验证。评估了近端策略优化(PPO)结合循环网络、无效动作掩码及探索策略的配置。实验表明,动作掩码显著提升策略性能,而仅依赖惩罚反馈外加正向奖励塑造,进一步带来收益。
原文摘要 · Abstract (English)
Routing vessels through narrow and dynamic waterways is challenging due to changing environmental conditions and operational constraints. Existing vessel-routing studies typically fail to generalize across multiple origin-destination pairs and do not exploit large-scale, data-driven traffic graphs. In this paper, we propose a reinforcement learning solution for big maritime data that can learn to find a route across multiple origin-destination pairs while adapting to different hexagonal grid resolutions. Agents learn to select direction and speed under continuous observations in a multi-discrete action space. A reward function balances fuel efficiency, travel time, wind resistance, and route diversity, using an Automatic Identification System (AIS)-derived traffic graph with ERA5 wind fields. The approach is demonstrated in the Gulf of St. Lawrence, one of the largest estuaries in the world. We evaluate configurations that combine Proximal Policy Optimization with recurrent networks, invalid-action masking, and exploration strategies. Our experiments demonstrate that action masking yields a clear improvement in policy performance and that supplementing penalty-only feedback with positive shaping rewards produces additional gains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。