用认知模型让大模型更真实模拟人类说服博弈中的决策行为
Using Cognitive Models to Improve Language Model Simulation of Human Persuasion Games

- 用方程驱动提示词引导大模型模仿贝叶斯更新等认知机制
- 大模型能准确模拟多种决策模式,小模型经强化学习后信念误差降26.5%
- 适合研究人类决策、训练多样化智能体或构建更真实评估环境
人们在策略互动中决策方式各异:有的按贝叶斯方式更新信念,有的则受动机性推理等偏差影响。尽管大语言模型常使用模拟人类进行安全评估与训练,但难以覆盖人类行为的多样性。我们提出‘方程转行为提示’方法,利用认知科学中的数学决策模型指导大模型行为。在基于法律决策的说服博弈上测试发现,大模型可通过提示拟合贝叶斯更新、仿射扭曲、动机性更新及Grether的α-β模型,而小模型无法做到。但通过强化学习训练小模型遵循数学规则(方程转行为强化学习),其在分布外参数设置下的信念误差降低26.5%。此外,训练小模型考虑不同类型的决策者,使其平均信念变化比仅训练贝叶斯型提升2.5%–12%,即使面对GPT-5-mini也有效。该方法可提升训练与评估中的人类模拟真实性,并为复杂人类决策模型研究提供新路径。
原文摘要 · Abstract (English)
People make decisions differently in strategic interactions. Some update beliefs like a Bayesian; others exhibit biases like motivated reasoning. Although creators of large language models use simulated humans for safety evaluations and training, they often fail to cover this breadth of human behavior. We argue that cognitive science and economics provide a convenient tool for doing so, making use of mathematical models of human decision-making. We propose an approach that we call Equation-to-Behavior Prompting for guiding large language models to match cognitive models, and evaluate this approach on persuasion games based on legal decision-making. We find that large models can approximate equation-based specifications -- Bayesian updating, affine distortion, motivated updating, and Grether's $α$-$β$ model -- using prompting, but small models fail to do so. However, training small models with reinforcement learning to adhere to mathematical rules, Equation-to-Behavior RL, reduces belief error by 26.5% in out-of-distribution parameterizations. We show that these simulations can help create diverse training environments; training small models to consider different kinds of decision-makers improves average belief change by 2.5%--12% over Bayesian-only training, even when persuading GPT-5-mini. Our work could improve human simulations for training and evaluation in increasingly realistic settings, and could also enable novel research into more complicated mathematical models of human decision-making.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。