SocialJax加速多智能体社会困境实验,效率提升50倍以上。
SocialJax: An Evaluation Suite for Multi-agent Reinforcement Learning in Sequential Social Dilemmas
- 基于JAX构建高效仿真环境,显著提升训练速度。
- 相比Melting Pot RLlib基线,实测速度提升超50倍。
- 适合关注多智能体协作与博弈的算法研究者。
序列式社会困境是多智能体强化学习中的关键挑战,需设计能准确反映个体与集体利益冲突的环境。现有基准如Melting Pot虽提供新社交伙伴泛化评估协议,但传统环境运行强化学习算法需大量计算资源。本文提出SocialJax,一个基于JAX实现的序列社会困境环境与算法套件。JAX作为高性能数值计算库,极大提升了运算效率。实验表明,SocialJax训练流水线相较Melting Pot RLlib基线实现至少50倍的实时性能提升。我们验证了基线算法在该环境中的有效性,并通过谢林图(Schelling diagrams)确认环境具备真实社会困境特性,确保其能准确捕捉社会博弈动态。
原文摘要 · Abstract (English)
Sequential social dilemmas pose a significant challenge in the field of multi-agent reinforcement learning (MARL), requiring environments that accurately reflect the tension between individual and collective interests. Previous benchmarks and environments, such as Melting Pot, provide an evaluation protocol that measures generalization to new social partners in various test scenarios. However, running reinforcement learning algorithms in traditional environments requires substantial computational resources. In this paper, we introduce SocialJax, a suite of sequential social dilemma environments and algorithms implemented in JAX. JAX is a high-performance numerical computing library for Python that enables significant improvements in operational efficiency. Our experiments demonstrate that the SocialJax training pipeline achieves at least 50\texttimes{} speed-up in real-time performance compared to Melting Pot RLlib baselines. Additionally, we validate the effectiveness of baseline algorithms within SocialJax environments. Finally, we use Schelling diagrams to verify the social dilemma properties of these environments, ensuring that they accurately capture the dynamics of social dilemmas.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。