GPT-4.1在模拟赌博中表现出与人类行为一致的风险偏好,且受设定身份影响。
Persona-Conditioned Risk Behavior in Large Language Models: A Simulated Gambling Study with GPT-4.1
- 给GPT-4.1分配贫富身份后,让其在三种老虎机环境下自主决策。
- 穷人身份模型平均玩37.4轮,富人仅1.1轮,差异极显著(p<2.2e-16)。
- 模型行为符合前景理论,但情绪标签不影响决策,适合研究认知偏见隐含性。
大型语言模型(LLMs)越来越多地被部署为不确定环境下的自主代理。然而,它们在这些环境中表现出的行为是源于原则性认知模式,还是仅仅表面化的提示模仿仍不清楚。本文通过一项受控实验,将GPT-4.1分配为三个社会经济身份(富有、中等收入、贫困),并置于包含三种机器配置的结构化老虎机环境中:公平(50%)、低偏置(35%)和连败加成(动态概率随连续失败上升)。每种条件进行50次独立迭代,共记录6,950次决策。结果发现,模型在未被指令的情况下重现了卡尼曼与特沃斯基前景理论预测的关键行为特征。贫困身份平均每会话玩37.4轮(标准差15.5),而富有身份仅1.1轮(标准差0.31),差异高度显著(Kruskal-Wallis H=393.5,p<2.2e-16)。不同身份间风险评分效应量巨大(贫困对富有Cohen's d=4.15)。情绪标签呈现为事后注释而非决策驱动(卡方=3205.4,Cramer's V=0.39),且跨轮次信念更新可忽略不计(贫困身份斯皮尔曼等级相关rho=0.032,p=0.016)。这些发现对大模型代理设计、可解释性研究以及经典认知经济偏见是否在大规模预训练模型中隐含编码具有重要启示。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly deployed as autonomous agents in uncertain, sequential decision-making contexts. Yet it remains poorly understood whether the behaviors they exhibit in such environments reflect principled cognitive patterns or simply surface-level prompt mimicry. This paper presents a controlled experiment in which GPT-4.1 was assigned one of three socioeconomic personas (Rich, Middle-income, and Poor) and placed in a structured slot-machine environment with three distinct machine configurations: Fair (50%), Biased Low (35%), and Streak (dynamic probability increasing after consecutive losses). Across 50 independent iterations per condition and 6,950 recorded decisions, we find that the model reproduces key behavioral signatures predicted by Kahneman and Tversky's Prospect Theory without being instructed to do so. The Poor persona played a mean of 37.4 rounds per session (SD=15.5) compared to 1.1 rounds for the Rich persona (SD=0.31), a difference that is highly significant (Kruskal-Wallis H=393.5, p<2.2e-16). Risk scores by persona show large effect sizes (Cohen's d=4.15 for Poor vs Rich). Emotional labels appear to function as post-hoc annotations rather than decision drivers (chi-square=3205.4, Cramer's V=0.39), and belief-updating across rounds is negligible (Spearman rho=0.032 for Poor persona, p=0.016). These findings carry implications for LLM agent design, interpretability research, and the broader question of whether classical cognitive economic biases are implicitly encoded in large-scale pretrained language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。