让大模型通过专家与自身体验反思,提升星际争霸游戏策略能力
Reflection of Episodes: Learning to Play Game from Expert and Self Experiences
- 结合专家经验与自我反思,动态优化决策过程
- 在文本版星际争霸难级下击败机器人,胜率显著提升
- 适合研究强化学习与大模型交互的学者参考
星际争霸Ⅱ是一个复杂且动态的实时战略游戏环境,非常适合人工智能与强化学习研究。为解决大语言模型在复杂环境中通过自我反思学习的问题,我们提出基于专家经验与自身体验的回合反思框架(ROE)。该框架首先通过关键帧选择方法提取游戏中关键信息,再结合专家经验与自身体验做出决策。游戏结束后,对过往经历进行反思以生成新的自身体验。实验表明,该方法在文本版星际争霸Ⅱ的“非常困难”难度下击败了内置机器人,并详细分析了大模型在游戏过程中的行为数据,验证了其有效性。
原文摘要 · Abstract (English)
StarCraft II is a complex and dynamic real-time strategy (RTS) game environment, which is very suitable for artificial intelligence and reinforcement learning research. To address the problem of Large Language Model(LLM) learning in complex environments through self-reflection, we propose a Reflection of Episodes(ROE) framework based on expert experience and self-experience. This framework first obtains key information in the game through a keyframe selection method, then makes decisions based on expert experience and self-experience. After a game is completed, it reflects on the previous experience to obtain new self-experience. Finally, in the experiment, our method beat the robot under the Very Hard difficulty in TextStarCraft II. We analyze the data of the LLM in the process of the game in detail, verified its effectiveness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。