arXiv:2411.04915cs.LGcs.AI2024-11被引 4

对比强化学习算法在内河航运中的抗干扰能力,发现SAC更稳定。

Evaluating Robustness of Reinforcement Learning Algorithms for Autonomous Shipping

  • 用模型无关的SAC算法在模拟器中实现有效航行策略。
  • SAC在未训练过的港口环境中仍能成功导航,鲁棒性优于MuZero。
  • 适合关注自主船舶安全与复杂环境适应性的研究者。

近年来,由于提升海事效率和安全的潜力,自主航运受到广泛关注。人工智能等先进技术可应对当前自主航运面临的导航与操作挑战。特别是内河航道运输(IWT)存在航道拥挤、环境多变等独特难题。在动态环境下,自主航运解决方案的可靠性与鲁棒性对保障安全运营至关重要。本文评估了基准深度强化学习(RL)算法在内河航运模拟器中的表现,及其生成有效运动规划策略的能力。结果表明,无模型方法可在模拟器中获得良好策略,成功导航训练中未遇的港口环境。特别地,我们发现软演员-评论家(SAC)相比最先进的基于模型的算法MuZero,具有更强的环境扰动鲁棒性。本研究为开发可泛化至多种船型及复杂港口与内河场景的鲁棒强化学习框架迈出关键一步。

原文摘要 · Abstract (English)

Recently, there has been growing interest in autonomous shipping due to its potential to improve maritime efficiency and safety. The use of advanced technologies, such as artificial intelligence, can address the current navigational and operational challenges in autonomous shipping. In particular, inland waterway transport (IWT) presents a unique set of challenges, such as crowded waterways and variable environmental conditions. In such dynamic settings, the reliability and robustness of autonomous shipping solutions are critical factors for ensuring safe operations. This paper examines the robustness of benchmark deep reinforcement learning (RL) algorithms, implemented for IWT within an autonomous shipping simulator, and their ability to generate effective motion planning policies. We demonstrate that a model-free approach can achieve an adequate policy in the simulator, successfully navigating port environments never encountered during training. We focus particularly on Soft-Actor Critic (SAC), which we show to be inherently more robust to environmental disturbances compared to MuZero, a state-of-the-art model-based RL algorithm. In this paper, we take a significant step towards developing robust, applied RL frameworks that can be generalized to various vessel types and navigate complex port- and inland environments and scenarios.

强化学习自主航运鲁棒性SAC

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。