LLMs在策略博弈中表现不佳,因感知、信念与行动间存在断裂。
Why Do LLMs Struggle in Strategic Play? Broken Links Between Observations, Beliefs, and Actions

- 通过分析内部信念发现,模型对游戏状态的判断比口头表达更准但易受干扰。
- 多跳推理时信念准确率下降,且出现首因与近因偏差,违背贝叶斯一致性。
- 内部信念转为行动的能力弱于显式提示引导,部署前需加防护机制。
大型语言模型(LLMs)越来越多地被用于不完全信息下的策略决策任务,如谈判和政策制定。尽管在某些任务上表现优异,其失败模式仍不清晰。本文通过实验揭示了基于Llama 3.1、Qwen3和gpt-oss等开源模型的两大根本性机制缺陷。第一,观察-信念差距:模型对隐藏游戏状态的内部信念远比其口头报告更准确,但信念脆弱——多跳推理时准确率下降,表现出首因与近因偏差,并在长期交互中偏离贝叶斯一致性。第二,信念-行动差距:内部信念隐式转化为行动的能力弱于显式提示中的信念引导,但后者也未持续提升博弈收益。这些结果表明,剖析模型内部过程可暴露系统性弱点,在缺乏稳健保护机制前,应谨慎将LLMs应用于战略领域。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly tasked with strategic decision-making under incomplete information, such as in negotiation and policymaking. While LLMs can excel at many such tasks, they also fail in ways that are poorly understood. We shed light on these failures by uncovering two fundamental gaps in the internal mechanisms underlying the decision-making of LLMs in incomplete-information games, supported by experiments with open-weight models Llama 3.1, Qwen3, and gpt-oss. First, an observation-belief gap: LLMs encode internal beliefs about latent game states that are substantially more accurate than their own verbal reports, yet these beliefs are brittle. In particular, the belief accuracy degrades with multi-hop reasoning, exhibits primacy and recency biases, and drifts away from Bayesian coherence over extended interactions. Second, a belief-action gap: The implicit conversion of internal beliefs into actions is weaker than that of the beliefs externalized in the prompt, yet neither belief-conditioning consistently achieves higher game payoffs. These results show how analyzing LLMs' internal processes can expose systematic vulnerabilities that warrant caution before deploying LLMs in strategic domains without robust guardrails.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。