用扩散模型辅助多智能体追踪,解决水下机器人协同跟踪难题。
Diffusion-Guided Cooperative Policy Learning for Target Tracking Based on Underwater Mobile Agent Networks
- 双决策策略融合扩散生成与确定性控制,分离经验池减少干扰。
- 精选高质量轨迹反向扩散,引导策略聚焦有效动作区域。
- 行为克隆损失对齐探索与执行,抑制动态环境下的策略漂移。
多智能体强化学习(MARL)为自主水下航行器(AUV)网络的协同目标追踪提供了有前景的解决方案。然而,现有方法仍面临三大挑战:1)多智能体并发更新导致的策略非平稳性;2)探索过程中积累的经验质量异质性引起的低效学习;3)在动态水下扰动下,随机探索与确定性执行之间的策略漂移。为此,本文提出四层分层MARL架构,包含全局训练调度、多智能体协调、局部策略生成和实时动作执行。基于此,设计了监督扩散辅助的MARL(SDA-MARL)算法,包含三个紧密耦合机制:首先,双决策策略结合基于扩散的生成分支与深度确定性策略梯度(DDPG)分支,通过隔离经验池降低训练干扰;其次,监督样本选择机制识别高质量追踪转移,利用其动作指导逆向扩散,使生成策略聚焦有效动作空间;第三,行为克隆损失将扩散生成的动作迁移至确定性DDPG Actor,实现探索与执行对齐,抑制策略漂移。在六自由度水下环境中,针对多种AUV-目标配置的实验表明,SDA-MARL相比基准MARL方法具备更快收敛速度、更高追踪精度、更一致的多智能体速度以及更短的追踪路径。
原文摘要 · Abstract (English)
Multi-agent reinforcement learning (MARL) provides a promising solution for cooperative target tracking in networks of autonomous underwater vehicles (AUVs). However, existing methods still face three major challenges: 1) policy non-stationarity caused by concurrent updates among multiple agents; 2) inefficient policy learning caused by the heterogeneous quality of experiences accumulated during exploration; and 3) policy drift between stochastic exploration and deterministic execution under dynamic underwater disturbances. To address these challenges, this paper develops a four-layer hierarchical MARL architecture comprising global training scheduling, multi-agent coordination, local policy generation, and real-time action execution. Building on this architecture, we propose a Supervised Diffusion-Aided MARL (SDA-MARL) algorithm with three closely coupled mechanisms. First, a dual-decision policy integrates a diffusion-based generative branch with a Deep Deterministic Policy Gradient (DDPG) branch, while segregated experience pools reduce training interference between the two branches. Second, a supervised sample-selection mechanism identifies high-quality tracking transitions and uses their actions to guide reverse diffusion, enabling the generative policy to concentrate on effective regions of the action space. Third, a behavioral-cloning loss transfers diffusion-generated actions to the deterministic DDPG Actor, thereby aligning exploration with execution and suppressing policy drift. Experiments conducted in six-degree-of-freedom underwater environments across multiple AUV-target configurations show that SDA-MARL achieves faster convergence, higher tracking accuracy, more consistent inter-AUV velocities, and shorter tracking paths than the compared MARL methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。