persona如何影响AI在博弈中合作或背叛
How Personas Can Influence Agents to Play Split or Steal

- 用角色设定引导大模型在重复博弈中决策
- 74%回合达成双方合作,背叛率低于11%
- 亲社会人格最易合作,分析型人格更倾向欺骗
角色设定常用于引导大型语言模型代理,但在社交困境情境中其对策略行为的影响仍不明确。为此,我们研究了角色提示在迭代式‘分或抢’博弈中的作用,其中角色驱动的代理与由固定提示控制的虚拟人类(VH)互动。代理基于四种开源模型(Ministral 3:3b、phi4:14b、Gemma3:12b、Gemma4:e4b)在两个温度设置(0.3 和 0.7)及确定性决策(零温度)下生成,而虚拟人类则由 GPT 4.1 mini 驱动。共进行160轮实验,每轮15回合,使用欧洲葡萄牙语。结果显示,双方合作(Split)占主导地位(约74%回合),背叛(Steal)仅发生于少于11%的回合。模型选择显著影响行为:phi4 和 Ministral 3:3b 在不同温度下均保持稳定合作,而 Gemma3:12b 与 Gemma4:e4b 表现出更多样化的策略与结果。基于五大性格特质的分析表明,亲社会(Prosocial)与原则性(Principled)角色最一致地合作,分析型(Analytical)角色更倾向于欺骗虚拟人类。话题分析显示,友谊相关对话与分钱决策一致,金钱与复仇相关内容更常见于抢夺结果;情感标签多为中性或积极,对解释结果帮助有限。这些发现刻画了角色提示与模型差异在重复信任博弈中的交互机制,并为后续涉及人类参与者与具身虚拟人类的虚拟现实研究提供了基准。
原文摘要 · Abstract (English)
Personas are often employed to guide large language model agents, yet their effectiveness in shaping strategic behavior in social dilemma settings remains uncertain. To address this, we examined the impact of persona prompts in an iterated Split or Steal game where persona-driven agents interacted with a Virtual Human (VH) controlled by a fixed prompt. Agents were instantiated from four open models (Ministral 3:3b, phi4:14b, Gemma3:12b, and Gemma4:e4b) at two temperature settings (0.3 and 0.7) and deterministic decision with zero temperature, while the VH was powered by GPT 4.1 mini. Across 160 sessions of 15 rounds each conducted in European Portuguese, mutual Split outcomes dominated (roughly 74 percent of rounds), with exploitation occurring in fewer than 11 percent of rounds. Model choice significantly influenced behavior: phi4 and Ministral 3:3b remained consistently cooperative across temperatures, whereas Gemma3:12b and Gemma4:e4b exhibited more varied strategies and outcomes. Analyses based on Big Five personality traits indicated that Prosocial and Principled personas were most consistently cooperative, while Analytical personas were more likely to exploit the VH. Topic analysis revealed that friendship-related dialogue aligns with Split decisions, whereas money and vengeance-related content is more prevalent in Steal outcomes; sentiment labels were predominantly neutral or happy and provided limited additional explanatory value. These findings characterize the interaction between persona prompts and model differences in repeated trust games and serve as a baseline for planned virtual reality studies involving human participants interacting with an embodied VH.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。