异构智能体协作学习,显著加速训练并提升性能。
Group-Agent Reinforcement Learning with Heterogeneous Agents
- 设计异步知识共享机制,动态决定是否及如何借鉴他人动作选择。
- 96%的智能体实现加速,72%提速超100倍,41%用不到5%时间达成更高得分。
- 适用于多算法混合的复杂协作场景,适合分布式强化学习研究者。
群体强化学习(GARL)是一种新兴的学习范式,多个强化学习智能体以异步方式共同学习并共享知识,旨在提升个体学习性能。在更一般的异构设置下,不同智能体使用不同学习算法,本文提出新颖有效的群体学习机制,指导智能体判断是否以及如何从其他智能体的动作选择中学习,并允许采纳表现更优的策略与价值函数模型。我们在总计43种Atari 2600游戏中进行了广泛实验,结果显示,在129个被评估的智能体中,96%实现了学习速度提升,72%的学习速度提升超过100倍;约41%的智能体在不足单个智能体独立学习所需5%的时间步内,获得了更高的累积奖励分数。
原文摘要 · Abstract (English)
Group-agent reinforcement learning (GARL) is a newly arising learning scenario, where multiple reinforcement learning agents study together in a group, sharing knowledge in an asynchronous fashion. The goal is to improve the learning performance of each individual agent. Under a more general heterogeneous setting where different agents learn using different algorithms, we advance GARL by designing novel and effective group-learning mechanisms. They guide the agents on whether and how to learn from action choices from the others, and allow the agents to adopt available policy and value function models sent by another agent if they perform better. We have conducted extensive experiments on a total of 43 different Atari 2600 games to demonstrate the superior performance of the proposed method. After the group learning, among the 129 agents examined, 96% are able to achieve a learning speed-up, and 72% are able to learn over 100 times faster. Also, around 41% of those agents have achieved a higher accumulated reward score by learning in less than 5% of the time steps required by a single agent when learning on its own.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。