大模型规模提升能显著改善多数社会模拟,但低资源领域效果有限。
Will Scaling Improve Social Simulation with LLMs?
- 基于计算量与参数规模的缩放定律,评估三类社会模拟任务表现。
- 多数观点与行为模拟随规模提升快速改进,尤其在英语语料丰富的群体中。
- 对认知偏差与长周期预测任务,规模效应弱,需专项研究支持。
大型语言模型(LLM)的社会模拟具有潜力,但目前仍缺乏足够的真实性。本文通过缩放定律研究计算规模、通用能力基准与三类代表性社会模拟任务(观点建模、行为模拟、长期预测)之间的关系。我们使用85个基于Qwen3架构、在DCLM网页文本语料上预训练的Transformer LLM,在固定计算预算(10¹⁸至10²⁰ FLOPs)下进行实验。随后评估了35个更大更强大的开源模型(最大70B参数),发现可从损失值预测下游准确率。结果表明,多数行为与观点模拟任务将随规模迅速提升,尤其在英语网络语料中代表充分的群体。长期预测与未充分代表的观点则增长较慢,尤其当其与通用知识和推理基准(如MMLU)相关性较低时。在行为模拟中,规模提升未能改善模型对人类认知偏差(如风险规避)或启发式策略(如从相关任务学习奖励)的校准能力,即使微调后,0.5B到8B参数的模型性能也无明显提升。综合来看,规模可提升大多数场景下的社会模拟质量,但存在例外,低资源领域改进不可靠。
原文摘要 · Abstract (English)
Large Language Model (LLM) social simulations are a promising research method, but they are not yet faithful enough to be adopted widely. In this work, we investigate whether the current scaling paradigm in language modeling is likely to close these gaps, or whether simulation fidelity is orthogonal to general capabilities and therefore deserving of more research attention. We use scaling laws to study the relationship between LLMs' compute scale, general capability benchmarks, and the fidelity of social simulation in three representative sub-domains: opinion modeling, behavioral simulation, and longitudinal forecasting. Surprisingly, we discover strong compute scaling in all three settings, using a suite of 85 transformer LLMs with the Qwen3 architecture pre-trained on the DCLM web text corpus under fixed-compute budgets from $10^{18}$ to $10^{20}$ FLOPs. Then we evaluate 35 larger and more capable open-weight models up to 70B parameters, allowing us to predict downstream accuracy from loss. This reveals that the majority of behavioral and opinion simulation tasks will rapidly improve with scale, particularly when they involve populations that are well-represented in English web corpora. Longitudinal forecasting and underrepresented opinions scale more slowly, especially when they are less correlated with general knowledge and reasoning benchmarks like MMLU. In behavior simulation, scaling fails to improve model calibration with human cognitive biases like risk aversion, as well as human heuristics like learning correlated rewards from related tasks. On these tasks, even fine-tuned models fail to noticeably scale up performance from 0.5B to 8B parameters. Taken together, we conclude that scale will improve social simulations in most settings, but outliers exist, and improvements will be less reliable in low-resource domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。