区分模拟与预测,基础模型更适合作个人意见生成。
Emulate or Estimate? The Divergent Strengths of Base and Post-Trained Language Models for Opinion Simulation

- 区分模拟个体响应与预测总体分布两类任务
- 基础模型生成结果更贴近真实人群分布与人口结构
- 后训练模型更适合直接预测总体分布,适合统计任务
大型语言模型被越来越多地用于模拟人类观点,但现有研究结果矛盾:部分研究发现模型与人类调查数据高度一致,另一些则报告人格坍塌和弱人口敏感性。我们提出,这些冲突源于混淆了两种不同任务。第一种是模拟(emulation),即模型生成个体回答,使整体分布趋近真实;第二种是估计(estimation),即模型直接预测总体分布。在Pew美国趋势面板数据集上评估六组匹配的基础模型与后训练模型,发现基础模型作为模拟器表现更优:生成的响应分布更接近人类真实分布,且更好保留人口结构特征。后训练模型则普遍在分布预测任务中表现更强,能更准确预测总体分布。因此,选择模型应基于任务需求:若需生成文本,则选基础模型;若只需预测分布,则选后训练模型。
原文摘要 · Abstract (English)
Large language models are increasingly used to simulate human opinions, but prior work reports conflicting results: some studies find promising alignment with human survey data, while others find persona collapse and weak demographic sensitivity. We propose that much of this conflict stems from conflating two distinct tasks. We call the first task emulation, in which models generate individual responses that aggregate into a population distribution. We call the second task estimation, in which models directly predict the population distribution. Evaluating six matched base and post-trained models on the Pew American Trends Panel, we find that base models are the stronger emulators: they produce response distributions closer to human ground truth and better preserve demographic structure. Post-trained models are generally the stronger estimators, producing more accurate distributional predictions when asked directly. We argue that model selection for human simulation should be guided by whether the task requires generating text or predicting distributions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。