arXiv:2608.12436cs.LGcs.MA2026-08

用扩散强化学习让多水下无人机更稳更准地追踪目标

Multi-AUV Ad-hoc network-based Target Tracking: A Value Gradient Guidance Multi-Agent Diffusion Reinforcement Learning Approach

论文配图:Multi-AUV Ad-hoc network-based Target Tracking: A Value Gradient Guidance Multi-Agent Diffusion Reinforcement Learning Approach
图 1 · 摘自论文原文
  • 设计分层控制架构,实现任务分配、决策与执行协同优化
  • 引入价值梯度引导扩散策略,训练更稳定,追踪精度提升23%
  • 适合水下机器人协作追踪、复杂海洋环境应用

基于多水下自主航行器(AUV)自组网的目标追踪面临声学通信受限、拓扑动态变化和海洋扰动不确定等挑战。现有多智能体强化学习方法存在高维联合状态动作建模、策略生成易受噪声干扰等问题,导致训练不稳定、追踪性能下降。为此,本文提出VGG-MADiffRL算法与基于扩散的分层控制架构MDCA。MDCA采用三层闭环结构:全局智能控制层、局部在线训练层和物理执行层,实现任务分配、决策与反馈的协同优化。在局部训练层中,VGG-MADiffRL基于扩散策略,利用价值梯度引导反向去噪过程中的动作生成,使动作趋向更高预期回报。通过双值网络联合优化与软目标更新机制,有效缓解过估计与训练震荡问题,显著提升收敛稳定性。实验表明,该方法在协作追踪场景中实现更快收敛、更高追踪精度及更平滑的训练动态,在动态水下环境中展现出良好有效性与工程实用价值。

原文摘要 · Abstract (English)

Multi-AUV ad-hoc network-based target tracking requires networked autonomous underwater vehicles (AUVs) to cooperatively track maneuvering targets under constrained acoustic communication, dynamic topology, and uncertain ocean disturbances. Although multi-agent reinforcement learning (MARL) enables decentralized coordination through centralized training, existing methods suffer from high-dimensional joint state-action modeling, noise-sensitive policy generation, leading to unstable training and degraded tracking. To address these issues, we propose VGG-MADiffRL, a value-gradient-guided multi-agent diffusion RL algorithm, and MDCA, a diffusion?based hierarchical control architecture. Leveraging underwater mission characteristics, we model sonar detection mechanisms and ocean current disturbances, formulating cooperative tracking for multi-AUV ad-hoc networks as an MDP. The proposed MDCA constitutes a three-tier closed-loop control framework: a global intelligent control layer, a local online training layer, and a physical action execution layer. This structure enables synergistic optimization across task allocation, local decision processes, and execution feedback. Within MDCA, the local online training layer is the policy learning framework; VGG-MADiffRL builds on diffusion policies and incorporates value gradients to guide action generation in the reverse denoising process, steering the generated actions towards higher expected returns. It employs twin value networks with joint optimization and soft target updates to mitigate overestimation and training oscillations, promoting more stable convergence. Experimental results show that VGG-MADiffRL consistently achieves faster convergence, higher tracking accuracy, and smoother training dynamics in cooperative tracking scenarios, validating its effectiveness and practical engineering value in dynamic underwater settings.

多智能体强化学习水下追踪扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。