用少量成本训练出超越人类顶尖水平的国际象棋类游戏AI
Superhuman AI for Stratego Using Self-Play Reinforcement Learning and Test-Time Search
- 通过自对弈强化学习与测试时搜索提升决策能力
- 在数千美元成本下达到远超人类顶尖水平的表现
- 为隐含信息下的战略博弈提供低成本高效解决方案
少数经典游戏被视为人工智能的重要基准,其训练成本可达数百万美元。其中,战略棋类游戏Stratego因极度复杂的隐藏信息环境,长期难以达到顶级人类水平。本文提出一种通用方法,结合自对弈强化学习与测试时搜索,在仅需数千美元成本下,实现远超人类顶尖水平的表现。该成果标志着Stratego领域性能与成本效率的重大突破。
原文摘要 · Abstract (English)
Few classical games have been regarded as such significant benchmarks of artificial intelligence as to have justified training costs in the millions of dollars. Among these, Stratego -- a board wargame exemplifying the challenge of strategic decision making under massive amounts of hidden information -- stands apart as a case where such efforts failed to produce performance at the level of top humans. This work establishes a step change in both performance and cost for Stratego, showing that it is now possible not only to reach the level of top humans, but to achieve vastly superhuman level -- and that doing so requires not an industrial budget, but merely a few thousand dollars. We achieved this result by developing general approaches for self-play reinforcement learning and test-time search under imperfect information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。