首个标准化基准测试大模型模拟人类行为的能力。
SimBench: Benchmarking the Ability of Large Language Models to Simulate Human Behaviors
- 构建涵盖20个数据集的统一评估框架,覆盖道德决策到经济选择等任务。
- 当前最佳模型模拟准确率仅40.80/100,性能随模型规模对数增长。
- 模拟能力与知识推理强相关,但对特定人群模拟效果差。
大语言模型(LLM)对人类行为的模拟若能忠实反映真实行为,将可能彻底改变社会科学。然而,现有评估方法零散,依赖定制任务与指标,导致结果难以比较。为此,我们提出SimBench,首个大规模、标准化的基准,旨在推动可复现的LLM模拟科学研究。通过整合20个多样化数据集,覆盖从道德决策到经济选择的任务,并基于大规模全球参与者群体,为探究模型在何时、如何及为何成功或失败提供基础。结果显示,当前最优模型的模拟准确率为40.80/100,性能随模型规模呈对数线性增长,但不受推理时计算量影响。我们发现存在模拟与对齐的权衡:指令微调提升低熵(共识)问题表现,但损害高熵(多样)问题表现。模型在模拟特定人口群体时尤为困难。最后,模拟能力与知识密集型推理高度相关(MMLU-Pro,r = 0.939)。通过使进展可度量,我们致力于加速更精准的LLM模拟器发展。
原文摘要 · Abstract (English)
Large language model (LLM) simulations of human behavior have the potential to revolutionize the social and behavioral sciences, if and only if they faithfully reflect real human behaviors. Current evaluations of simulation fidelity are fragmented, based on bespoke tasks and metrics, creating a patchwork of incomparable results. To address this, we introduce SimBench, the first large-scale, standardized benchmark for a robust, reproducible science of LLM simulation. By unifying 20 diverse datasets covering tasks from moral decision-making to economic choice across a large global participant pool, SimBench provides the necessary foundation to ask fundamental questions about when, how, and why LLM simulations succeed or fail. We show that the best LLMs today achieve meaningful but modest simulation fidelity (score: 40.80/100), with performance scaling log-linearly with model size but not with increased inference-time compute. We discover an alignment-simulation tradeoff: instruction tuning improves performance on low-entropy (consensus) questions but degrades it on high-entropy (diverse) ones. Models particularly struggle when simulating specific demographic groups. Finally, we demonstrate that simulation ability correlates most strongly with knowledge-intensive reasoning (MMLU-Pro, r = 0.939). By making progress measurable, we aim to accelerate the development of more faithful LLM simulators.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。