arXiv:2601.18580cs.LG2026-01

用多样化智能体并行探索,加速强化学习训练并发现多样解。

K-Myriad: Jump-starting reinforcement learning with unsupervised parallel agents

  • 通过多智能体并行提升状态熵,实现无监督多样化探索。
  • 在高维连续控制任务中,生成多种不同策略,提升训练效率。
  • 适合需要高效探索与多样解的强化学习场景。

强化学习中的并行化通常用于加速单一策略的训练,多个工作者从相同的采样分布中收集经验。这种设计限制了并行化的潜力,忽略了多样化探索策略的优势。我们提出 K-Myriad,一种可扩展且无监督的方法,通过群体并行策略最大化集体状态熵。通过培育一组专业化探索策略,K-Myriad 为强化学习提供稳健的初始化,显著提升训练效率并发现异质性解决方案。在大规模并行化的高维连续控制任务上,实验表明 K-Myriad 能学习到一系列不同的策略,验证其在集体探索中的有效性,并为新型并行化策略铺平道路。

原文摘要 · Abstract (English)

Parallelization in Reinforcement Learning is typically employed to speed up the training of a single policy, where multiple workers collect experience from an identical sampling distribution. This common design limits the potential of parallelization by neglecting the advantages of diverse exploration strategies. We propose K-Myriad, a scalable and unsupervised method that maximizes the collective state entropy induced by a population of parallel policies. By cultivating a portfolio of specialized exploration strategies, K-Myriad provides a robust initialization for Reinforcement Learning, leading to both higher training efficiency and the discovery of heterogeneous solutions. Experiments on high-dimensional continuous control tasks, with large-scale parallelization, demonstrate that K-Myriad can learn a broad set of distinct policies, highlighting its effectiveness for collective exploration and paving the way towards novel parallelization strategies.

强化学习并行化探索策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。