arXiv:2509.02579cs.LGcs.AI2025-09被引 3

用隐变量建模提升无人机协同反盗猎的决策能力

Latent Variable Modeling in Multi-Agent Reinforcement Learning via Expectation-Maximization for UAV-Based Wildlife Protection

  • 通过期望最大化算法构建隐变量模型,捕捉环境与多机互动的未知因素
  • 10架无人机在伊朗豹保护区仿真中,检测准确率和策略收敛性显著优于PPO/DDPG
  • 适合关注复杂环境下的分布式智能决策与生态保护应用的研究者

保护濒危野生动物免受非法盗猎威胁,在广阔且部分可观测的环境中尤为严峻,实时响应至关重要。本文提出一种基于期望最大化(EM)的多智能体强化学习(MARL)隐变量建模方法,用于无人机(UAV)协同野生动物保护。通过隐变量建模隐藏环境因素与智能体间动态,增强不确定性下的探索与协作能力。我们在一个包含10架无人机的定制仿真环境中评估该方法,任务为巡逻濒危伊朗豹的保护栖息地。大量实验结果表明,相比标准算法如近端策略优化(PPO)和深度确定性策略梯度(DDPG),本方法在检测准确率、适应性和策略收敛性方面表现更优。研究证实,将EM推断与MARL结合可显著提升复杂高风险保护场景中的去中心化决策能力。完整实现、仿真环境及训练脚本已公开于GitHub。

原文摘要 · Abstract (English)

Protecting endangered wildlife from illegal poaching presents a critical challenge, particularly in vast and partially observable environments where real-time response is essential. This paper introduces a novel Expectation-Maximization (EM) based latent variable modeling approach in the context of Multi-Agent Reinforcement Learning (MARL) for Unmanned Aerial Vehicle (UAV) coordination in wildlife protection. By modeling hidden environmental factors and inter-agent dynamics through latent variables, our method enhances exploration and coordination under uncertainty.We implement and evaluate our EM-MARL framework using a custom simulation involving 10 UAVs tasked with patrolling protected habitats of the endangered Iranian leopard. Extensive experimental results demonstrate superior performance in detection accuracy, adaptability, and policy convergence when compared to standard algorithms such as Proximal Policy Optimization (PPO) and Deep Deterministic Policy Gradient (DDPG). Our findings underscore the potential of combining EM inference with MARL to improve decentralized decisionmaking in complex, high-stakes conservation scenarios. The full implementation, simulation environment, and training scripts are publicly available on GitHub.

多智能体强化学习无人机协同隐变量建模生态保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。