让不同结构的智能体安全协作找目标,靠通信正则化避免碰撞。
Safe Heterogeneous Multi-Agent RL with Communication Regularization for Coordinated Target Acquisition
- 用图注意力网络融合感知与邻居通信信息,实现上下文决策。
- 引入安全过滤器和正交通信向量设计,提升任务稳定性与安全性。
- 适合多智能体协同搜索、无人机编队等需避障的任务场景。
本文提出一种去中心化的多智能体强化学习框架,使结构异构的智能体团队能在部分可观测、通信受限且动态交互的环境中协同发现并获取随机分布的目标。每个智能体的策略采用多智能体近端策略优化(MAD-PPO)算法训练,并使用图注意力网络编码器,融合模拟测距数据与邻近智能体间交换的通信嵌入,实现基于局部感知与关系信息的上下文感知决策。特别地,该工作构建了一个统一框架,结合基于图的通信机制与轨迹感知的安全性保障,通过安全过滤器实现动态避障。系统采用结构化奖励函数,鼓励有效目标发现、碰撞规避及智能体间通信向量的去相关性,通过促进信息正交性实现。全面消融实验验证了奖励函数的有效性。仿真结果表明,该框架可实现安全稳定的任务执行,证实其有效性。
原文摘要 · Abstract (English)
This paper introduces a decentralized multi-agent reinforcement learning framework enabling structurally heterogeneous teams of agents to jointly discover and acquire randomly located targets in environments characterized by partial observability, communication constraints, and dynamic interactions. Each agent's policy is trained with the Multi-Agent Proximal Policy Optimization algorithm and employs a Graph Attention Network encoder that integrates simulated range-sensing data with communication embeddings exchanged among neighboring agents, enabling context-aware decision-making from both local sensing and relational information. In particular, this work introduces a unified framework that integrates graph-based communication and trajectory-aware safety through safety filters. The architecture is supported by a structured reward formulation designed to encourage effective target discovery and acquisition, collision avoidance, and de-correlation between the agents' communication vectors by promoting informational orthogonality. The effectiveness of the proposed reward function is demonstrated through a comprehensive ablation study. Moreover, simulation results demonstrate safe and stable task execution, confirming the framework's effectiveness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。