用卡牌游戏分析大模型对规则协同的理解能力
Rule Synergy Analysis using LLMs: State of the Art and Implications
- 构建卡牌协同数据集,评估大模型对规则互动的推理能力
- 模型能准确识别无协同卡牌,但对正向和负向协同识别率低
- 揭示时序、状态定义等关键错误类型,指导未来模型改进
大型语言模型(LLMs)在逻辑推理、数学等多个领域表现出色。本文研究大模型在动态环境(如卡牌游戏)中理解与推理复杂规则交互的能力。我们基于游戏《消灭幽灵》(Slay the Spire)构建了一个卡牌协同数据集,其中卡牌组合被分类为正向、负向或中性协同。评估显示,尽管大模型在识别非协同组合方面表现良好,但在检测正向协同及尤其负向协同方面存在明显不足。我们归纳了常见错误类型,包括时机判断偏差、游戏状态定义不清以及规则遵循错误。研究结果为提升模型预测规则及其交互效果的能力指明了未来方向。
原文摘要 · Abstract (English)
Large language models (LLMs) have demonstrated strong performance across a variety of domains, including logical reasoning, mathematics, and more. In this paper, we investigate how well LLMs understand and reason about complex rule interactions in dynamic environments, such as card games. We introduce a dataset of card synergies from the game Slay the Spire, where pairs of cards are classified based on their positive, negative, or neutral interactions. Our evaluation shows that while LLMs excel at identifying non-synergistic pairs, they struggle with detecting positive and, particularly, negative synergies. We categorize common error types, including issues with timing, defining game states, and following game rules. Our findings suggest directions for future research to improve model performance in predicting the effect of rules and their interactions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。