arXiv:2507.16306cs.MAcs.RO2025-07中稿 · IEEE MRS 2025被引 1

多智能体协作追踪动态目标,提升监测效率与鲁棒性。

COMPASS: Cooperative Multi-Agent Persistent Monitoring using Spatio-Temporal Attention Network

  • 基于时空注意力网络实现去中心化协同决策
  • 在动态目标场景下显著降低不确定性并提升覆盖率
  • 适合灾害响应、环境监测等实时追踪任务

持续监测动态目标在灾难救援、环境感知和野生动物保护等实际应用中至关重要,需在不确定性下由移动智能体持续获取信息。本文提出COMPASS,一种多智能体强化学习框架,使去中心化智能体能高效持久地监测多个移动目标。将环境建模为图结构,节点代表空间位置,边表示拓扑邻近关系,支持智能体基于结构化布局推理并适时重访高信息区域。每个智能体独立依据共享的时空注意力网络做出动作选择,该网络整合历史观测与空间上下文。目标动态采用高斯过程(GPs)建模,支持合理的信念更新并实现不确定性感知规划。使用集中式价值估计和去中心化策略执行,在自适应奖励设置下训练。大量实验表明,COMPASS在不确定性降低、目标覆盖和协调效率方面均优于强基线,在动态多目标场景中表现一致优异。

原文摘要 · Abstract (English)

Persistent monitoring of dynamic targets is essential in real-world applications such as disaster response, environmental sensing, and wildlife conservation, where mobile agents must continuously gather information under uncertainty. We propose COMPASS, a multi-agent reinforcement learning (MARL) framework that enables decentralized agents to persistently monitor multiple moving targets efficiently. We model the environment as a graph, where nodes represent spatial locations and edges capture topological proximity, allowing agents to reason over structured layouts and revisit informative regions as needed. Each agent independently selects actions based on a shared spatio-temporal attention network that we design to integrate historical observations and spatial context. We model target dynamics using Gaussian Processes (GPs), which support principled belief updates and enable uncertainty-aware planning. We train COMPASS using centralized value estimation and decentralized policy execution under an adaptive reward setting. Our extensive experiments demonstrate that COMPASS consistently outperforms strong baselines in uncertainty reduction, target coverage, and coordination efficiency across dynamic multi-target scenarios.

多智能体强化学习目标追踪时空建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。