arXiv:2604.04157cs.AI2026-04被引 1

LLM poker agents靠记忆自发形成心理模型,展现类心智理论行为。

Readable Minds: Emergent Theory-of-Mind-Like Behavior in LLM Poker Agents

  • 用持久记忆的LLM代理打德州扑克,通过互动自动生成对手心理模型。
  • 有记忆的代理达到三级以上预测与递归建模,无记忆者始终停留在零级。
  • 模型以自然语言表达,可读性强,适合研究AI社会智能与人类认知机制。

心智理论(ToM)——对他人心理状态的建模能力——是人类社会认知的核心。以往关于大语言模型(LLMs)是否具备ToM的研究仅限于静态情景,未能检验其在动态交互中是否可自发产生。本文报告,在持续进行德州扑克对局的自主LLM代理中,只有配备持久记忆的代理才会逐步发展出复杂对手模型。在2×2因子设计中,交叉测试记忆(存在/缺失)与领域知识(存在/缺失),每组重复5次(共20次实验,约6000次代理-手牌观察),结果表明:记忆是生成类ToM行为的必要且充分条件(Cliff's delta = 1.0, p = 0.008)。有记忆的代理达到ToM Level 3–5(从预测到递归建模),无记忆者则始终维持在Level 0。基于对手模型的战略欺骗仅出现在有记忆条件下(Fisher精确检验,p < 0.001)。领域知识虽不决定类ToM行为的出现,但能提升其应用精度:无扑克知识的代理仍能达到同等水平的建模,但欺骗更不精准(p = 0.004)。具备ToM的代理偏离游戏理论最优策略(TAG遵守率67% vs. 79%,delta = -1.0, p = 0.008),转而针对性地针对特定对手,模仿专家人类玩家行为。所有心理模型均以自然语言表达,完全可读,为理解人工智能社会认知提供透明窗口。跨模型验证使用GPT-4o,加权Cohen's kappa = 0.81(几乎完全一致)。这些发现表明,功能性类心智理论行为可仅由交互动态自发产生,无需显式训练或提示,对理解人工社会智能与生物社会认知具有重要意义。

原文摘要 · Abstract (English)

Theory of Mind (ToM) -- the ability to model others' mental states -- is fundamental to human social cognition. Whether large language models (LLMs) can develop ToM has been tested exclusively through static vignettes, leaving open whether ToM-like reasoning can emerge through dynamic interaction. Here we report that autonomous LLM agents playing extended sessions of Texas Hold'em poker progressively develop sophisticated opponent models, but only when equipped with persistent memory. In a 2x2 factorial design crossing memory (present/absent) with domain knowledge (present/absent), each with five replications (N = 20 experiments, ~6,000 agent-hand observations), we find that memory is both necessary and sufficient for ToM-like behavior emergence (Cliff's delta = 1.0, p = 0.008). Agents with memory reach ToM Level 3-5 (predictive to recursive modeling), while agents without memory remain at Level 0 across all replications. Strategic deception grounded in opponent models occurs exclusively in memory-equipped conditions (Fisher's exact p < 0.001). Domain expertise does not gate ToM-like behavior emergence but enhances its application: agents without poker knowledge develop equivalent ToM levels but less precise deception (p = 0.004). Agents with ToM deviate from game-theoretically optimal play (67% vs. 79% TAG adherence, delta = -1.0, p = 0.008) to exploit specific opponents, mirroring expert human play. All mental models are expressed in natural language and directly readable, providing a transparent window into AI social cognition. Cross-model validation with GPT-4o yields weighted Cohen's kappa = 0.81 (almost perfect agreement). These findings demonstrate that functional ToM-like behavior can emerge from interaction dynamics alone, without explicit training or prompting, with implications for understanding artificial social intelligence and biological social cognition.

心智理论博弈智能可解释性LLM行为

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。