arXiv:2604.19404cs.ROcs.AI2026-04

用Mamba提升水下机器人协同追捕的长期决策与稳定性

M$^{2}$GRPO: Mamba-based Multi-Agent Group Relative Policy Optimization for Biomimetic Underwater Robots Pursuit

论文配图:M$^{2}$GRPO: Mamba-based Multi-Agent Group Relative Policy Optimization for Biomimetic Underwater Robots Pursuit
图 1 · 摘自论文原文
  • 基于Mamba的策略网络捕捉长时间依赖和机器人间关系
  • 组内奖励归一化使追捕成功率提升18.6%,资源需求降低40%
  • 适合需要稳定协作的仿生水下机器人系统研发者

仿生水下机器人协同追捕面临长周期决策、部分可观测性与多机协同的挑战,传统策略学习方法难以兼顾表达能力与稳定性。本文提出基于Mamba的多智能体组相对策略优化(M²GRPO)框架,在集中训练、分散执行(CTDE)范式下,融合选择性状态空间Mamba策略与组相对策略优化。Mamba策略利用观测历史捕捉长时序依赖,通过注意力机制编码智能体间交互,采用归一化高斯采样生成有界连续动作。为改善信用分配且不牺牲稳定性,通过组内归一化奖励计算组相对优势,并以多智能体扩展的GRPO进行优化,显著降低训练资源需求,实现稳定可扩展的策略更新。大规模仿真与真实泳池实验表明,该方法在不同团队规模与逃逸策略下,持续优于MAPPO与循环基线,在追捕成功率和捕获效率上表现更优。

原文摘要 · Abstract (English)

Traditional policy learning methods in cooperative pursuit face fundamental challenges in biomimetic underwater robots, where long-horizon decision making, partial observability, and inter-robot coordination require both expressiveness and stability. To address these issues, a novel framework called Mamba-based multi-agent group relative policy optimization (M$^{2}$GRPO) is proposed, which integrates a selective state-space Mamba policy with group-relative policy optimization under the centralized-training and decentralized-execution (CTDE) paradigm. Specifically, the Mamba-based policy leverages observation history to capture long-horizon temporal dependencies and exploits attention-based relational features to encode inter-agent interactions, producing bounded continuous actions through normalized Gaussian sampling. To further improve credit assignment without sacrificing stability, the group-relative advantages are obtained by normalizing rewards across agents within each episode and optimized through a multi-agent extension of GRPO, significantly reducing the demand for training resources while enabling stable and scalable policy updates. Extensive simulations and real-world pool experiments across team scales and evader strategies demonstrate that M$^{2}$GRPO consistently outperforms MAPPO and recurrent baselines in both pursuit success rate and capture efficiency. Overall, the proposed framework provides a practical and scalable solution for cooperative underwater pursuit with biomimetic robot systems.

多智能体仿生机器人强化学习Mamba

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。