用对话模型复现人类决策偏差,揭示其在压力下的行为动态。
Emulating Aggregate Human Choice Behavior and Biases with GPT Conversational Agents
- 将经典决策实验转为对话场景,模拟认知负荷影响
- GPT-4/5 在1100人实验中精准复现人类偏差模式
- 适合研究人机交互中的偏见建模与系统设计
认知偏差常影响人类决策。尽管大语言模型(LLMs)已被证明能再现已知偏差,但更关键的问题是:它们能否在上下文因素(如认知负荷)作用下,预测并模拟个体层面的偏差动态?我们把三个经典的决策场景改编为对话形式,并开展人类实验(N=1100)。参与者通过简单或复杂对话与聊天机器人互动完成决策。结果揭示了显著的偏差现象。为评估LLM在相似交互条件下如何模仿人类决策,我们使用参与者的年龄、性别等人口统计信息及对话文本,基于GPT-4和GPT-5构建模拟环境。结果显示,两种模型均以高精度重现了人类偏差。不同模型在对人类行为的拟合程度上存在明显差异,这对设计和评估交互式环境中具备偏见感知能力的AI系统具有重要意义。
原文摘要 · Abstract (English)
Cognitive biases often shape human decisions. While large language models (LLMs) have been shown to reproduce well-known biases, a more critical question is whether LLMs can predict biases at the individual level and emulate the dynamics of biased human behavior when contextual factors, such as cognitive load, interact with these biases. We adapted three well-established decision scenarios into a conversational setting and conducted a human experiment (N=1100). Participants engaged with a chatbot that facilitates decision-making through simple or complex dialogues. Results revealed robust biases. To evaluate how LLMs emulate human decision-making under similar interactive conditions, we used participant demographics and dialogue transcripts to simulate these conditions with LLMs based on GPT-4 and GPT-5. The LLMs reproduced human biases with precision. We found notable differences between models in how they aligned human behavior. This has important implications for designing and evaluating adaptive, bias-aware LLM-based AI systems in interactive contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。