人类在红白棋比赛中胜过大模型,因更擅长策略推理。
Not Yet: Humans Outperform LLMs in a Colonel Blotto Tournament

- 人类采用更合理的中间层级分配策略,优于模型的刻板打法。
- 人类在三轮比赛中持续胜出,尤其在高阶推理时表现更好。
- 适合关注博弈策略与人类决策的研究者阅读。
大型语言模型(LLMs)的兴起促使经济学家研究人类与模型在策略环境中的行为差异。我们组织了多轮红白棋博弈锦标赛:第一轮有200多名人类参赛者相互对战;第二轮邀请多个主流大模型提交策略;第三轮将模型数量与人类人数匹配。结果发现,人类更常使用经过良好校准的中等复杂度分配启发式方法,优于模型所提交的简单、刻板策略。战略复杂度是成功关键,仅当达到必要推理深度时才有效,过低或过高均无优势。人类中,理工科背景者在首轮表现略优。令人意外的是,人类在不同对手组合的赛事间几乎不调整策略,表明其决策主要基于游戏规则而非对手身份,将模型视为与人类同等的对手。
原文摘要 · Abstract (English)
The emergence of large language models (LLMs) has spurred economists to study how humans and LLMs behave in strategic settings. We organized a series of round-robin tournaments in the Colonel Blotto game. This game attracts game theorists' attention due to high-dimensional action space and the absence of pure strategy Nash equilibria. In the first tournament, more than 200 human participants competed against one another. In the second tournament, several popular LLMs were invited to submit strategies. In the third tournament, we matched the number of LLM strategies to the number submitted by humans. We find that humans more often employ better-calibrated intermediate-level allocation heuristics and outperform the simpler, more stereotyped strategies submitted by LLMs. Strategic sophistication is key to success if and only if the necessary level of reasoning depth is reached, while lower and higher levels of reasoning offer no clear advantage over the primitive strategies. Among humans, field of study weakly predicts success: participants with STEM backgrounds perform better in the first tournament. Surprisingly, humans almost do not adjust their strategies across tournaments with different sets of opponents. This result suggests that humans base their choices primarily on the game's rules rather than on the identity of their opponents, treating LLMs much like human competitors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。