结合在线离线方法,用超算快速学会下西洋棋,水平逼近顶尖人类和计算机玩家。
Learning and Improving Backgammon Strategy
- 融合离线训练与在线蒙特卡洛模拟,实现并行化策略优化
- 在短时间内达到与顶尖人类和计算机玩家相当的对弈水平
- 适合研究强化学习与大规模并行计算的学者和爱好者
本文提出一种新方法,结合在线与离线学习特性,利用并行超级计算机的强大算力,在短时间内高效学习西洋棋价值函数。离线方法包括神经网络并行训练与TD(λ)强化学习;本文引入大规模并行的蒙特卡洛“回溯”技术作为在线策略改进手段,将计算资源集中于博弈树搜索中的决策点,进一步提升价值函数估计。实验表明,该方法在短时间内实现了与当前顶尖人类及计算机西洋棋选手水平相当甚至可能更优的对弈表现。
原文摘要 · Abstract (English)
A novel approach to learning is presented, combining features of on-line and off-line methods to achieve considerable performance in the task of learning a backgammon value function in a process that exploits the processing power of parallel supercomputers. The off-line methods comprise a set of techniques for parallelizing neural network training and $TD(λ)$ reinforcement learning; here Monte-Carlo ``Rollouts'' are introduced as a massively parallel on-line policy improvement technique which applies resources to the decision points encountered during the search of the game tree to further augment the learned value function estimate. A level of play roughly as good as, or possibly better than, the current champion human and computer backgammon players has been achieved in a short period of learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。