用多智能体强化学习追踪多个污染源,仅需探索1.29%环境即可定位。
Multi-source Plume Tracing via Multi-Agent Reinforcement Learning
- 将污染源追踪建模为部分可观测马尔可夫博弈,利用历史观测与动作序列增强决策能力
- 在仿真环境中仅探索1.29%的区域即成功定位多个污染源,显著优于传统方法
- 适用于复杂湍流场景下的应急监测,适合无人机群协同环境感知任务
工业灾难如1984年博帕尔事件和2015年阿尔索·坎昌天然气泄漏凸显了快速可靠污染羽流追踪算法对公共健康与环境安全的重要性。传统梯度法或仿生方法在真实湍流条件下常失效。本文提出一种多智能体强化学习(MARL)算法,利用小型无人飞行器集群(sUAS)定位多个空气污染物源。该方法将问题建模为部分可观测马尔可夫博弈(POMG),采用基于LSTM的双深度循环Q网络(ADDRQN),使用完整的动作-观测历史序列来近似潜在状态。不同于以往工作,我们基于高斯羽流模型(GPM)构建通用仿真环境,包含三维空间、传感器噪声、多智能体交互及多污染源等真实要素。通过将动作历史作为输入,提升了模型在复杂部分可观测环境中的适应性。大量仿真表明,本算法显著优于传统方法:模型仅需探索1.29%的环境即可成功定位污染源。
原文摘要 · Abstract (English)
Industrial catastrophes like the Bhopal disaster (1984) and the Aliso Canyon gas leak (2015) demonstrate the urgent need for rapid and reliable plume tracing algorithms to protect public health and the environment. Traditional methods, such as gradient-based or biologically inspired approaches, often fail in realistic, turbulent conditions. To address these challenges, we present a Multi-Agent Reinforcement Learning (MARL) algorithm designed for localizing multiple airborne pollution sources using a swarm of small uncrewed aerial systems (sUAS). Our method models the problem as a Partially Observable Markov Game (POMG), employing a Long Short-Term Memory (LSTM)-based Action-specific Double Deep Recurrent Q-Network (ADDRQN) that uses full sequences of historical action-observation pairs, effectively approximating latent states. Unlike prior work, we use a general-purpose simulation environment based on the Gaussian Plume Model (GPM), incorporating realistic elements such as a three-dimensional environment, sensor noise, multiple interacting agents, and multiple plume sources. The incorporation of action histories as part of the inputs further enhances the adaptability of our model in complex, partially observable environments. Extensive simulations show that our algorithm significantly outperforms conventional approaches. Specifically, our model allows agents to explore only 1.29\% of the environment to successfully locate pollution sources.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。