用分层强化学习优化救援队协同找伤员,提升响应效率。
Factorized Deep Q-Network for Cooperative Multi-Agent Reinforcement Learning in Victim Tagging
- 设计分层DQN模型,让多救援员自主决策并局部通信。
- 小规模场景下模型比传统规则快23%,复杂场景仍逊于启发式方法。
- 验证了局部通信+动态重规划在不确定性环境中的优势,适合应急响应研究者。
大规模伤亡事件(MCIs)因复杂性和不确定性日益成为关注焦点,伤员标记环节需快速完成,以支持后续紧迫的救援行动。本文针对多智能体协同伤员标记问题,提出数学建模以最小化标记总时长。评估了五种分布式启发式策略,涵盖全局与局部通信能力下的不同情境。进一步对比基于因子化深度Q网络(FDQN)的多智能体强化学习策略与基线启发式方法。大量仿真实验表明:具备局部通信的策略在自适应标记中更高效,尤其选择最近伤员并支持重规划时表现突出。在小规模场景中,FDQN优于启发式方法;但在复杂场景下,启发式策略更具优势。实验覆盖多样复杂度,探索了多智能体强化学习在现实应用中的极限,揭示关键洞见。
原文摘要 · Abstract (English)
Mass casualty incidents (MCIs) are a growing concern, characterized by complexity and uncertainty that demand adaptive decision-making strategies. The victim tagging step in the emergency medical response must be completed quickly and is crucial for providing information to guide subsequent time-constrained response actions. In this paper, we present a mathematical formulation of multi-agent victim tagging to minimize the time it takes for responders to tag all victims. Five distributed heuristics are formulated and evaluated with simulation experiments. The heuristics considered are on-the go, practical solutions that represent varying levels of situational uncertainty in the form of global or local communication capabilities, showcasing practical constraints. We further investigate the performance of a multi-agent reinforcement learning (MARL) strategy, factorized deep Q-network (FDQN), to minimize victim tagging time as compared to baseline heuristics. Extensive simulations demonstrate that between the heuristics, methods with local communication are more efficient for adaptive victim tagging, specifically choosing the nearest victim with the option to replan. Analyzing all experiments, we find that our FDQN approach outperforms heuristics in smaller-scale scenarios, while heuristics excel in more complex scenarios. Our experiments contain diverse complexities that explore the upper limits of MARL capabilities for real-world applications and reveal key insights.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。