arXiv:2505.06997cs.AI2025-05被引 1

多智能体强化学习优化救援中空地人协同感知任务分配

A Multi-Agent Reinforcement Learning Approach for Cooperative Air-Ground-Human Crowdsensing in Emergency Rescue

  • 设计'硬合作'策略,地面车优先为低电量无人机充电
  • 在时间约束下使任务完成率提升18.42%
  • 适合应急救援场景的实时协同决策系统

移动众包正从传统人类中心模式演进,融合无人机(UAV)、无人车(UGV)等异构实体。在复杂环境、通信受限、部分可观测的应急救援场景中,优化多类智能体的任务分配至关重要。本文针对空-地-人协同感知任务分配(HECTA)问题,提出一种‘硬合作’策略:地面车在执行感知任务的同时,优先为低电量无人机充电。目标是最大化任务完成率(TCR),在严格时间约束下实现。将该问题建模为分布式部分可观测马尔可夫决策过程(Dec-POMDP),以应对不确定性下的序列决策。为此提出HECTA4ER算法,基于中央训练、分散执行架构,包含专用特征提取模块、利用历史动作-观测信息的隐状态机制,以及融合全局与局部信息的混合网络,有效解决部分可观测性挑战。理论分析证明算法收敛性。大量仿真表明,相比基线方法,平均提升18.42%任务完成率;真实案例研究验证其在动态场景中的有效性与鲁棒性,具备实际应用潜力。

原文摘要 · Abstract (English)

Mobile crowdsensing is evolving beyond traditional human-centric models by integrating heterogeneous entities like unmanned aerial vehicles (UAVs) and unmanned ground vehicles (UGVs). Optimizing task allocation among these diverse agents is critical, particularly in challenging emergency rescue scenarios characterized by complex environments, limited communication, and partial observability. This paper tackles the Heterogeneous-Entity Collaborative-Sensing Task Allocation (HECTA) problem specifically for emergency rescue, considering humans, UAVs, and UGVs. We introduce a novel ``Hard-Cooperative'' policy where UGVs prioritize recharging low-battery UAVs, alongside performing their sensing tasks. The primary objective is maximizing the task completion rate (TCR) under strict time constraints. We rigorously formulate this NP-hard problem as a decentralized partially observable Markov decision process (Dec-POMDP) to effectively handle sequential decision-making under uncertainty. To solve this, we propose HECTA4ER, a novel multi-agent reinforcement learning algorithm built upon a Centralized Training with Decentralized Execution architecture. HECTA4ER incorporates tailored designs, including specialized modules for complex feature extraction, utilization of action-observation history via hidden states, and a mixing network integrating global and local information, specifically addressing the challenges of partial observability. Furthermore, theoretical analysis confirms the algorithm's convergence properties. Extensive simulations demonstrate that HECTA4ER significantly outperforms baseline algorithms, achieving an average 18.42% increase in TCR. Crucially, a real-world case study validates the algorithm's effectiveness and robustness in dynamic sensing scenarios, highlighting its strong potential for practical application in emergency response.

多智能体强化学习应急救援协同感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。