多智能体动态分组协作,高效分配异构资源。
Decentralized Reinforcement Learning for Multi-Agent Multi-Resource Allocation via Dynamic Cluster Agreements
- 基于动态聚类的分组机制,让智能体自组织成临时小队。
- 在20个智能体、5种资源下仍保持稳定收益与高效协调。
- 适合大规模分布式系统资源调度,尤其资源可释放场景。
本文针对多智能体在去中心化环境下分配异构资源的挑战,提出Liquid-Graph-Time Clustering-IPPO(LGTC-IPPO)方法。该方法在独立近端策略优化(IPPO)基础上,引入动态聚类共识机制,使智能体能根据资源需求自适应形成并调整局部子团队。此去中心化协调策略降低了对全局信息的依赖,提升了可扩展性。我们在不同团队规模和资源分布条件下,将LGTC-IPPO与标准多智能体强化学习基线及集中式专家方案进行对比。实验表明,该方法在20个智能体、5类资源设置下仍实现更稳定的奖励、更强的协同能力,并在资源种类或智能体数量增加时表现出鲁棒性能。此外,动态聚类机制还支持资源释放场景下的高效再分配。
原文摘要 · Abstract (English)
This paper addresses the challenge of allocating heterogeneous resources among multiple agents in a decentralized manner. Our proposed method, Liquid-Graph-Time Clustering-IPPO, builds upon Independent Proximal Policy Optimization (IPPO) by integrating dynamic cluster consensus, a mechanism that allows agents to form and adapt local sub-teams based on resource demands. This decentralized coordination strategy reduces reliance on global information and enhances scalability. We evaluate LGTC-IPPO against standard multi-agent reinforcement learning baselines and a centralized expert solution across a range of team sizes and resource distributions. Experimental results demonstrate that LGTC-IPPO achieves more stable rewards, better coordination, and robust performance even as the number of agents or resource types increases. Additionally, we illustrate how dynamic clustering enables agents to reallocate resources efficiently also for scenarios with discharging resources.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。