通过状态新颖性动态重用数据,提升多智能体强化学习效率与多样性
Novelty-Guided Data Reuse for Efficient and Diversified Multi-Agent Reinforcement Learning
- 基于随机网络蒸馏评估状态新颖性,动态分配更新机会
- 在Google Research Football和星海争霸微操任务中显著提升性能
- 适合追求高效探索与多样化策略的多智能体系统研究者
近年来,深度多智能体强化学习(MARL)在解决复杂协作任务方面展现出巨大潜力,推动了人工智能在协同环境中的边界。然而,这些系统的效率常因样本利用不足和学习策略缺乏多样性而受限。为提升MARL性能,我们提出一种新型样本重用方法,根据观察的新颖性动态调整策略更新。具体而言,采用随机网络蒸馏(RND)网络衡量每个智能体当前状态的新颖性,并依据数据独特性赋予额外的样本更新机会。该方法名为多智能体新颖性引导样本重用(MANGER),有效提升样本效率,促进探索并生成多样化的智能体行为。实验验证表明,在Google Research Football和超难星海争霸微管理任务等复杂协作场景中,MANGER显著提升了MARL的有效性。
原文摘要 · Abstract (English)
Recently, deep Multi-Agent Reinforcement Learning (MARL) has demonstrated its potential to tackle complex cooperative tasks, pushing the boundaries of AI in collaborative environments. However, the efficiency of these systems is often compromised by inadequate sample utilization and a lack of diversity in learning strategies. To enhance MARL performance, we introduce a novel sample reuse approach that dynamically adjusts policy updates based on observation novelty. Specifically, we employ a Random Network Distillation (RND) network to gauge the novelty of each agent's current state, assigning additional sample update opportunities based on the uniqueness of the data. We name our method Multi-Agent Novelty-GuidEd sample Reuse (MANGER). This method increases sample efficiency and promotes exploration and diverse agent behaviors. Our evaluations confirm substantial improvements in MARL effectiveness in complex cooperative scenarios such as Google Research Football and super-hard StarCraft II micromanagement tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。