多智能体在未知环境搜索中,用信息素逆向引导提升动态目标定位效率。
PILOC: A Pheromone Inverse Guidance Mechanism and Local-Communication Framework for Dynamic Target Search of Multi-Agent in Unknown Environments
- 基于深度强化学习,将信息素机制嵌入观测空间实现间接协作。
- 局部通信与信息素引导结合,使搜索效率提升37%以上。
- 适合通信受限、目标动态变化的搜救场景,如灾难救援。
多智能体搜索与救援(MASAR)在灾害响应、探索和侦察中至关重要。然而,动态且未知的环境因目标不可预测和环境不确定性带来巨大挑战。为此,我们提出PILOC框架,无需全局先验知识,依赖局部感知与通信。其引入信息素逆向引导机制,实现高效协同与动态目标定位。通过局部通信促进去中心化协作,显著降低对全局通道的依赖。不同于传统启发式方法,该信息素机制嵌入深度强化学习(DRL)的观测空间,支持基于环境线索的间接协作。我们将此策略集成至DRL多智能体架构,并进行广泛实验。结果表明,局部通信与基于信息素的引导结合,显著提升搜索效率、适应性与系统鲁棒性。相较于现有方法,PILOC在动态与通信受限场景下表现更优,为未来MASAR应用提供新方向。
原文摘要 · Abstract (English)
Multi-Agent Search and Rescue (MASAR) plays a vital role in disaster response, exploration, and reconnaissance. However, dynamic and unknown environments pose significant challenges due to target unpredictability and environmental uncertainty. To tackle these issues, we propose PILOC, a framework that operates without global prior knowledge, leveraging local perception and communication. It introduces a pheromone inverse guidance mechanism to enable efficient coordination and dynamic target localization. PILOC promotes decentralized cooperation through local communication, significantly reducing reliance on global channels. Unlike conventional heuristics, the pheromone mechanism is embedded into the observation space of Deep Reinforcement Learning (DRL), supporting indirect agent coordination based on environmental cues. We further integrate this strategy into a DRL-based multi-agent architecture and conduct extensive experiments. Results show that combining local communication with pheromone-based guidance significantly boosts search efficiency, adaptability, and system robustness. Compared to existing methods, PILOC performs better under dynamic and communication-constrained scenarios, offering promising directions for future MASAR applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。