LLMs show人性倾向,但互动推理不稳定。
Social preferences with unstable interactive reasoning: Large language models in economic trust games
- 让LLM参与信任博弈,测试其社会偏好与推理能力
- 无提示下仍表现信任与互惠,但多人回合时差异显著
- 角色设定影响远超模型或游戏类型,适合研究人机交互
尽管大语言模型(LLMs)在理解人类语言方面表现出色,本研究探讨其如何将这种理解转化为体现真实人际互动本质的经济信任博弈场景。实验中将ChatGPT-4、Claude和Bard置于需平衡自利与信任、互惠的博弈情境中,以揭示其社会偏好与交互推理能力。结果显示,即使未被引导扮演特定角色,这些模型也偏离纯自利行为,展现出信任与互惠倾向。在最简单的单轮互动中,它们模拟了人类玩家初始阶段的信任行为;而在涉及信任回报或多轮互动的情境中,决策受社会偏好与交互推理共同影响,差异更为明显。当被要求扮演自私或非自私角色时,模型响应变化显著,影响程度超过模型间或游戏类型差异。其中,以非自私或中立身份回应的ChatGPT-4表现出最高水平的信任与互惠,甚至超越人类、Claude和Bard;而Claude和Bard的表现则有时高于、有时低于人类。当被赋予自私角色时,所有模型的信任与互惠水平均低于人类。对对手行为或游戏机制变化的交互推理在各模型中呈现随机性,缺乏稳定可复现特征,尽管在某些角色设定下ChatGPT-4略有改善。
原文摘要 · Abstract (English)
While large language models (LLMs) have demonstrated remarkable capabilities in understanding human languages, this study explores how they translate this understanding into social exchange contexts that capture certain essences of real world human interactions. Three LLMs - ChatGPT-4, Claude, and Bard - were placed in economic trust games where players balance self-interest with trust and reciprocity, making decisions that reveal their social preferences and interactive reasoning abilities. Our study shows that LLMs deviate from pure self-interest and exhibit trust and reciprocity even without being prompted to adopt a specific persona. In the simplest one-shot interaction, LLMs emulated how human players place trust at the beginning of such a game. Larger human-machine divergences emerged in scenarios involving trust repayment or multi-round interactions, where decisions were influenced by both social preferences and interactive reasoning. LLMs responses varied significantly when prompted to adopt personas like selfish or unselfish players, with the impact outweighing differences between models or game types. Response of ChatGPT-4, in an unselfish or neutral persona, resembled the highest trust and reciprocity, surpassing humans, Claude, and Bard. Claude and Bard displayed trust and reciprocity levels that sometimes exceeded and sometimes fell below human choices. When given selfish personas, all LLMs showed lower trust and reciprocity than humans. Interactive reasoning to the actions of counterparts or changing game mechanics appeared to be random rather than stable, reproducible characteristics in the response of LLMs, though some improvements were observed when ChatGPT-4 responded in selfish or unselfish personas.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。