用522人真实数据测试大模型能否模拟经济行为,发现群体趋势可预测但个体难精准。
Pay What LLM Wants: Can LLM Simulate Economics Experiment with 522 Real-human Persona?
- 基于522名韩国参与者的真实人格数据,测试大模型在文化消费场景中的定价预测能力。
- 大模型在个体决策上表现不佳,但在群体行为趋势上具备合理预测力。
- 复杂提示方法不如简单提示,重构叙事或检索增强未带来明显提升。
近期大型语言模型(LLMs)在模拟人类行为方面取得进展,但多数研究依赖虚构人格而非真实人类数据。本研究通过522名韩国参与者在文化消费场景下的支付意愿(Pay-What-You-Want, PWYW)实验,评估大模型预测个体经济决策的能力。我们系统比较了三种先进多模态LLMs,使用详细的人格信息进行分析,探讨大模型是否能准确复现人类个体选择,以及人格注入方式对预测性能的影响。结果表明,尽管大模型在个体层面预测能力有限,但在群体行为趋势上表现出合理一致性。此外,常用提示技巧并不显著优于基础提示方法;个人叙事重建与检索增强生成也未带来明显优势。这些发现首次基于真实人类数据全面评估了大模型在经济行为模拟中的能力,为计算社会科学中基于人格的仿真提供了实证指导。
原文摘要 · Abstract (English)
Recent advances in Large Language Models (LLMs) have generated significant interest in their capacity to simulate human-like behaviors, yet most studies rely on fictional personas rather than actual human data. We address this limitation by evaluating LLMs' ability to predict individual economic decision-making using Pay-What-You-Want (PWYW) pricing experiments with real 522 human personas. Our study systematically compares three state-of-the-art multimodal LLMs using detailed persona information from 522 Korean participants in cultural consumption scenarios. We investigate whether LLMs can accurately replicate individual human choices and how persona injection methods affect prediction performance. Results reveal that while LLMs struggle with precise individual-level predictions, they demonstrate reasonable group-level behavioral tendencies. Also, we found that commonly adopted prompting techniques are not much better than naive prompting methods; reconstruction of personal narrative nor retrieval augmented generation have no significant gain against simple prompting method. We believe that these findings can provide the first comprehensive evaluation of LLMs' capabilities on simulating economic behavior using real human data, offering empirical guidance for persona-based simulation in computational social science.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。