arXiv:2502.06113cs.ROcs.SY2025-02

用生物启发策略加速水下多智能体强化学习,提升探索效率。

Towards Bio-inspired Heuristically Accelerated Reinforcement Learning for Adaptive Underwater Multi-Agents Behaviour

  • 引入粒子群优化启发式引导策略,提前聚焦高价值状态空间。
  • 在连续控制环境下,显著减少训练交互次数达到最优性能。
  • 适合需快速收敛的水下多机器人协同任务研究者使用。

本文研究复杂环境中自主多智能体系统的协同问题,目标是在覆盖规划任务中探测并识别感兴趣目标。该任务在太空与水下场景均具重要意义。传统上,覆盖规划被建模为马尔可夫决策过程,要求由异构自主水下航行器组成的群体对区域进行测绘与搜寻。该问题面临环境不确定性、通信约束及随时间变化的潜在风险等挑战。尽管多智能体强化学习(MARL)算法可通过深度神经网络解决高度非线性问题,并具备良好的可扩展性,但当前多数水下应用仍局限于仿真,因学习周期过长。为此,本文提出一种新型加速策略:将生物启发式方法——粒子群优化(PSO)引入训练过程,以指导策略在训练初期即聚焦于高价值的动作与状态空间,优化探索与利用的平衡。该方法应用于MSAC算法,在二维连续控制环境下完成覆盖任务评估,显著提升了收敛速度。

原文摘要 · Abstract (English)

This paper describes the problem of coordination of an autonomous Multi-Agent System which aims to solve the coverage planning problem in a complex environment. The considered applications are the detection and identification of objects of interest while covering an area. These tasks, which are highly relevant for space applications, are also of interest among various domains including the underwater context, which is the focus of this study. In this context, coverage planning is traditionally modelled as a Markov Decision Process where a coordinated MAS, a swarm of heterogeneous autonomous underwater vehicles, is required to survey an area and search for objects. This MDP is associated with several challenges: environment uncertainties, communication constraints, and an ensemble of hazards, including time-varying and unpredictable changes in the underwater environment. MARL algorithms can solve highly non-linear problems using deep neural networks and display great scalability against an increased number of agents. Nevertheless, most of the current results in the underwater domain are limited to simulation due to the high learning time of MARL algorithms. For this reason, a novel strategy is introduced to accelerate this convergence rate by incorporating biologically inspired heuristics to guide the policy during training. The PSO method, which is inspired by the behaviour of a group of animals, is selected as a heuristic. It allows the policy to explore the highest quality regions of the action and state spaces, from the beginning of the training, optimizing the exploration/exploitation trade-off. The resulting agent requires fewer interactions to reach optimal performance. The method is applied to the MSAC algorithm and evaluated for a 2D covering area mission in a continuous control environment.

多智能体强化学习水下机器人生物启发

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。