评估大模型模拟运营管理中人类行为的能力,发现其能复现关键结论但分布差异显著。
Predicting Effects, Missing Distributions: Evaluating LLMs as Human Behavior Simulators in Operations Management
- 用九项已有实验验证大模型能否复现人类决策模式。
- 大模型常能复制假设结果,但响应分布与真实数据差距大。
- 思维链提示和参数调优可有效缩小分布偏差,小模型也可媲美大模型。
大型语言模型(LLMs)在商业、经济和社会科学中被广泛用于模拟人类行为,为实验室实验、实地研究和调查提供低成本替代方案。本文基于九项已发表的行为运筹学实验,从两个维度评估LLM的性能:一是生成数据是否复现原始假设检验结果,二是其完整响应分布与人类数据的匹配度(以Wasserstein距离衡量)。结果显示,LLM通常能再现假设层面的效果,表明其可捕捉显著的决策偏差和行为规律;但其响应分布往往与真实人类数据存在显著偏离,尤其在离散程度上差异明显。此外,本文还考察了两种轻量级缓解策略:思维链提示(chain-of-thought prompting)和超参数调优。两者均能降低分布偏差,适当调优后,小型或开源模型甚至可达到或超越大型专有模型的表现。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used to simulate human behavior in business, economics, and the social sciences, offering a low-cost complement to laboratory experiments, field studies, and surveys. This paper evaluates how well LLMs replicate human behavior in operations management. Using nine published behavioral-operations experiments, we assess LLM performance along two dimensions: whether LLM-generated data reproduce the original hypothesis-test outcomes, and whether their full response distributions align with human data, measured by Wasserstein distance. We find that LLMs often replicate hypothesis-level effects, suggesting that they can capture salient decision biases and behavioral regularities. However, their response distributions frequently diverge from human data, even for strong proprietary models, with dispersion mismatch playing an important role. We also examine two lightweight mitigation strategies: chain-of-thought prompting and hyperparameter tuning. Both can reduce distributional misalignment, and appropriate tuning can sometimes allow smaller or open-source models to match or outperform larger proprietary systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。