评估大模型在决策中的行为真实性,发现其默认策略保守,难复制人类多样性。
Noise, Adaptation, and Strategy: Assessing LLM Fidelity in Decision-Making
- 通过三重干预框架(内在、指令、模仿)测试模型适应性
- 大模型默认趋同于稳定保守策略,与人类行为有显著差异
- 引入人类数据可缩小差距,但仍无法达到人类策略多样性
大型语言模型(LLMs)在社会科学模拟中应用日益广泛。尽管其在推理和优化任务上的表现已受广泛评估,但对其模拟人类决策变异性和适应性的能力关注不足。本文提出一种过程导向的评估框架,采用渐进式干预(内在性、指令、模仿),检验LLM代理在不同外部引导和人为噪声下的适应能力。在两个经典经济学任务——第二价格拍卖中的非理性行为、报童问题中的决策偏差上验证该框架,揭示了LLM与人类之间的行为差距。结果表明,大模型默认收敛于稳定且保守的策略,偏离真实人类行为;风险导向指令虽能可预测地影响行为,但未能复现人类多样性;通过上下文学习融入人类数据可缩小差距,但未达到人类受试者的策略变异性。这些发现凸显了行为保真度上的持续对齐缺口,提示未来评估应更关注过程层面的真实性。本研究提供了一种用于动态决策任务中评估LLMs的方法,为社会科学研究中的合成数据应用提供指导。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used in social science simulations. While their performance on reasoning and optimization tasks has been extensively evaluated, less attention has been paid to their ability to simulate human decision-making's variability and adaptability. We propose a process-oriented evaluation framework with progressive interventions (Intrinsicality, Instruction, and Imitation) to examine how LLM agents adapt under different levels of external guidance and human-derived noise. We validate the framework on two classic economics tasks, irrationality in the second-price auction and decision bias in the newsvendor problem, showing behavioral gaps between LLMs and humans. We find that LLMs, by default, converge on stable and conservative strategies that diverge from observed human behaviors. Risk-framed instructions impact LLM behavior predictably but do not replicate human-like diversity. Incorporating human data through in-context learning narrows the gap but fails to reach human subjects' strategic variability. These results highlight a persistent alignment gap in behavioral fidelity and suggest that future LLM evaluations should consider more process-level realism. We present a process-oriented approach for assessing LLMs in dynamic decision-making tasks, offering guidance for their application in synthetic data for social science research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。