arXiv:2510.03534cs.MAcs.LG2025-10中稿 · the 2026 IEEE Inte…

用多智能体强化学习高效监测杜罗河口的长期水体扩散

Long-Term Mapping of the Douro River Plume with Multi-Agent Reinforcement Learning

  • 中心协调器间歇通信,结合高斯过程与多头Q网络控制无人艇
  • 双倍无人艇数量可使续航翻倍,同时保持或提升监测精度
  • 策略跨季节泛化,适合长期动态环境监测

本文研究利用多个自主水下航行器(AUVs)对河流羽流进行长期(数日)监测,以杜罗河为例。提出一种节能且低通信量的多智能体强化学习方法,其中中心协调器间歇性地与各AUV通信,收集数据并下达指令。方法融合时空高斯过程回归(GPR)与多头Q网络控制器,分别调节每个AUV的方向和速度。基于Delft3D海洋模型的仿真表明,该方法在均方误差(MSE)和运行耐力方面均优于单/多智能体基准。在某些场景中,将无人艇数量加倍可使续航超过两倍,同时维持或提升准确性,凸显多智能体协作的优势。所学策略可在不同年份、季节的未见环境中泛化,展现出数据驱动的长期动态羽流监测前景。

原文摘要 · Abstract (English)

We study the problem of long-term (multiple days) mapping of a river plume using multiple autonomous underwater vehicles (AUVs), focusing on the Douro river representative use-case. We propose an energy - and communication - efficient multi-agent reinforcement learning approach in which a central coordinator intermittently communicates with the AUVs, collecting measurements and issuing commands. Our approach integrates spatiotemporal Gaussian process regression (GPR) with a multi-head Q-network controller that regulates direction and speed for each AUV. Simulations using the Delft3D ocean model demonstrate that our method consistently outperforms both single- and multi-agent benchmarks, with scaling the number of agents both improving mean squared error (MSE) and operational endurance. In some instances, our algorithm demonstrates that doubling the number of AUVs can more than double endurance while maintaining or improving accuracy, underscoring the benefits of multi-agent coordination. Our learned policies generalize across unseen seasonal regimes over different months and years, demonstrating promise for future developments of data-driven long-term monitoring of dynamic plume environments.

多智能体强化学习水体监测无人艇

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。