arXiv:2510.06970eess.SYcs.LG2025-10被引 2

用对抗场景训练船舶自主航行,让智能体更守规则。

Falsification-driven reinforcement learning for maritime motion planning

  • 通过生成违反航行规则的对抗场景来强化训练
  • 在双船开放海域实验中规则遵守率显著提升
  • 适合需要高安全性的海上自主系统研发者

遵守海上航行规则对自主船舶的安全运行至关重要,但训练强化学习(RL)智能体遵守这些规则极具挑战。智能体的行为由其经历的训练场景决定,而构建能体现海上航行复杂性的场景并不容易,仅靠真实数据也难以满足需求。为此,我们提出一种基于伪证驱动的强化学习方法,生成使被测船舶违反海上交通规则的对抗性训练场景,规则以信号时序逻辑规范表达。在双船开放海域导航实验中,该方法生成了更相关的训练场景,并实现了更一致的规则遵守表现。

原文摘要 · Abstract (English)

Compliance with maritime traffic rules is essential for the safe operation of autonomous vessels, yet training reinforcement learning (RL) agents to adhere to them is challenging. The behavior of RL agents is shaped by the training scenarios they encounter, but creating scenarios that capture the complexity of maritime navigation is non-trivial, and real-world data alone is insufficient. To address this, we propose a falsification-driven RL approach that generates adversarial training scenarios in which the vessel under test violates maritime traffic rules, which are expressed as signal temporal logic specifications. Our experiments on open-sea navigation with two vessels demonstrate that the proposed approach provides more relevant training scenarios and achieves more consistent rule compliance.

强化学习船舶导航规则合规

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。