用大模型模拟人类猜拳决策,揭示认知极限。
Understanding Human Limits in Pattern Recognition: A Computational Model of Sequential Reasoning in Rock, Paper, Scissors
- 用语言模型构建假设推理机制,模拟人类猜拳行为
- 在7个对手中,获自然语言提示后6个胜率超80%
- 揭示人类在复杂模式识别中的核心认知瓶颈
我们通过建模布罗克班克与沃尔(2024)实验中人类重复玩剪刀石头布的行为,探究人们如何从他人行为中预测模式以及计算能力的限制。面对策略复杂度不同的算法对手,人类能利用简单转移模式(如纸后必出石),但难以发现更复杂的序列依赖关系。为此,我们采用基于大语言模型的假设心智(HM)代理,模拟人类认知过程。结果显示,当应用于相同实验条件时,HM的表现模式与人类高度一致,在相似情境下成功或失败。通过一系列消融与增强实验,发现若提供对手策略的自然语言描述,HM可成功应对7个对手中的6个,胜率均高于80%,表明准确生成假设是主要认知瓶颈。进一步通过教育式干预系统性调整模型假设,发现其对对手行为的因果理解显著更新,显示基于模型的分析可为人类认知提出可验证假设。
原文摘要 · Abstract (English)
How do we predict others from patterns in their behavior and what are the computational constraints that limit this ability? We investigate these questions by modeling human behavior over repeated games of rock, paper, scissors from Brockbank & Vul (2024). Against algorithmic opponents that varied in strategic sophistication, people readily exploit simple transition patterns (e.g., consistently playing rock after paper) but struggle to detect more complex sequential dependencies. To understand the cognitive mechanisms underlying these abilities and their limitations, we deploy Hypothetical Minds (HM), a large language model-based agent that generates and tests hypotheses about opponent strategies, as a cognitive model of this behavior (Cross et al., 2024). We show that when applied to the same experimental conditions, HM closely mirrors human performance patterns, succeeding and failing in similar ways. To better understand the source of HM's failures and whether people might face similar cognitive bottlenecks in this context, we performed a series of ablations and augmentations targeting different components of the system. When provided with natural language descriptions of the opponents' strategies, HM successfully exploited 6/7 bot opponents with win rates >80% suggesting that accurate hypothesis generation is the primary cognitive bottleneck in this task. Further, by systematically manipulating the model's hypotheses through pedagogically-inspired interventions, we find that the model substantially updates its causal understanding of opponent behavior, revealing how model-based analyses can produce testable hypotheses about human cognition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。