arXiv:2608.12995cs.AIcs.MA2026-08中稿 · ICUS 2026

提出OGR-MARL框架,让异构无人船在受限港口高效协同追捕。

OGR-MARL: Option-Guided Residual Multi-Agent Reinforcement Learning for Heterogeneous USV Cooperative Pursuit in Constrained Port Waterways

论文配图:OGR-MARL: Option-Guided Residual Multi-Agent Reinforcement Learning for Heterogeneous USV Cooperative Pursuit in Constrained Port Waterways
图 1 · 摘自论文原文
  • 用选项引导+残差学习,让多智能体在规则约束下快速优化行为。
  • 在抽象港口场景中捕获率达75.0%,且无需重训练即可跨地图迁移。
  • 适合复杂水域的无人船协同控制,尤其对规则约束强的场景有效。

受限港口水域中的异构无人水面艇(USV)协同追捕任务,需在航行、交通和角色约束下实现追捕目标。本文提出一种不依赖特定强化学习算法的选项引导残差多智能体强化学习框架(OGR-MARL),融合共享追捕目标信念、角色条件化的选项目标、自适应规则惩罚与残差策略学习机制,使不同多智能体强化学习算法可基于规则引导行为学习修正动作,而非从零探索受限环境。将典型连续控制多智能体算法(如MADDPG、MATD3、MAPPO、MASAC)与OGR-MARL结合,构建OGR-MADDPG、OGR-MATD3、OGR-MAPPO、OGR-MASAC。在抽象的小芝门港口水域场景中,OGR-MASAC实现75.0%的捕获率,展现出良好任务有效性与规则合规性,并在异构协同方面优于其他方法。无需重训练,即可零样本迁移至基于QGIS/AIS信息的正式小芝门地图,取得良好效果,验证了OGR-MARL在更复杂港口场景中的泛化潜力。

原文摘要 · Abstract (English)

Heterogeneous USV cooperative pursuit in constrained port waterways requires evader interception under navigation, traffic, and role constraints. This paper proposes OGR-MARL, an option-guided residual multi-agent reinforcement learning framework that is decoupled from a specific MARL algorithm. OGR-MARL integrates shared evader belief, role-conditioned option targets, adaptive rule penalties, and residual policy learning, allowing different MARL algorithms to learn corrective actions on top of rule-guided behaviors rather than exploring constrained port environments from scratch. We instantiate OGR-MARL with representative continuous-control MARL backbones, including MADDPG, MATD3, MAPPO, and MASAC, yielding OGR-MADDPG, OGR-MATD3, OGR-MAPPO, and OGR-MASAC. Experiments in an abstract Xiazhimen port-waterway scenario show that the OGR-MASAC instantiation achieves a 75.0% capture rate, promising mission-effective rule compliance, and the best heterogeneous coordination among the tested methods. Without retraining, zero-shot transfer to a QGIS/AIS-informed Xiazhimen map achieves promising results, demonstrating the generalization potential of OGR-MARL in more complex port scenarios.

多智能体强化学习无人船协同追捕

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。