让大模型主动预演决策后果,提升游戏策略能力
What-If Analysis of Large Language Models: Explore the Game World Using Proactive Thinking
- 用自然语言构建可解释的世界模型,模拟动作后果
- 在王者荣耀中预测游戏状态变化准确率达74.2%,比基线高27%
- 适合研究游戏AI、具身智能与可解释决策的学者
大语言模型在高风险环境(如MOBA游戏)中决策表现不佳,主要源于缺乏主动推理和对复杂游戏动态的理解。为此,我们提出基于语言的世界模型——何若分析大模型(WiA-LLM),通过两阶段训练:先在人类推理轨迹上进行监督微调,再利用结果导向奖励进行强化学习,使模型能以自然语言预测候选动作引发的游戏状态演化,并提供文本解释。在《王者荣耀》(HoK)环境中,该模型对游戏状态变化的预测准确率达到74.2%,较基线模型提升27%。此外,其策略行为更接近专家玩家,展现出更强的前瞻性与专家级决策能力。
原文摘要 · Abstract (English)
LLMs struggle with decision-making in high-stakes environments like MOBA games, primarily due to a lack of proactive reasoning and limited understanding of complex game dynamics. To address this, we propose What-if Analysis LLM (WiA-LLM), a framework that trains an LLM as an explicit, language-based world model. Instead of representing the environment in latent vectors, WiA-LLM uses natural language to simulate how the game state evolves over time in response to candidate actions, and provides textual justifications for these predicted outcomes. WiA-LLM is trained in two stages: supervised fine-tuning on human-like reasoning traces, followed by reinforcement learning with outcome-based rewards based on the alignment between predicted and actual future states. In the Honor of Kings (HoK) environment, WiA-LLM attains 74.2\% accuracy (27\%$\uparrow$ vs. base model) in forecasting game-state changes. In addition, WiA-LLM demonstrate strategic behavior more closely aligned with expert players than purely reactive LLMs, indicating enhanced foresight and expert-like decision-making.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。