用单人游戏知识提升双人对战表现,训练更快更稳。
Enhancing Two-Player Performance Through Single-Player Knowledge Transfer: An Empirical Study on Atari 2600 Games
- 从单人版本迁移知识到双人环境,提升训练效率
- 在10个Atari 2600游戏中平均减少训练时间并提高得分
- 揭示内存复杂度与性能的关系,适合强化学习研究者
在双人对战环境中使用强化学习和自对弈训练面临挑战,主要源于环境复杂性及训练过程可能的不稳定性。本文提出,若强化学习算法能利用同一游戏的单人版本所学知识,可在双人游戏中实现更高效训练并获得更好性能。本研究在10个不同的Atari 2600游戏环境中进行了实证分析,以Atari 2600 RAM作为输入状态。结果表明,相比从零开始训练双人模型,通过单人知识迁移可显著缩短训练时间并提升平均总奖励。同时,本文还提出一种计算RAM复杂度的方法,并探讨其与模型表现之间的关联。
原文摘要 · Abstract (English)
Playing two-player games using reinforcement learning and self-play can be challenging due to the complexity of two-player environments and the possible instability in the training process. We propose that a reinforcement learning algorithm can train more efficiently and achieve improved performance in a two-player game if it leverages the knowledge from the single-player version of the same game. This study examines the proposed idea in ten different Atari 2600 environments using the Atari 2600 RAM as the input state. We discuss the advantages of using transfer learning from a single-player training process over training in a two-player setting from scratch, and demonstrate our results in a few measures such as training time and average total reward. We also discuss a method of calculating RAM complexity and its relationship to performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。