arXiv:2608.29455cs.CLcs.CY2026-08

大模型虽能模仿平均人类回答,却无法替代个体差异。

Item-Mean Surrogates: Why Richer Persona Data Fail to Improve LLMs as Human Surrogates

  • 用大模型模拟人类时,只预测平均值,忽略个体偏差。
  • 去除平均值后,模型解释个体差异仅3.05%,远低于人类重测信度53.6%。
  • 适合检验大模型是否真能替代真实人类的四种实证方法。

大模型被越来越多地用作人类替代者,常假设更丰富的个人资料可使其成为特定个体的代理或探索工具。我们在涵盖40余万参与者、6000多个调查项目和实验结果的四个数据集上测试这一假设。大模型在整体层面表现良好:其平均回答与人类平均回答高度一致。但这种成功主要源于对每个项目平均人类回答的预测。一旦移除项目平均值,大模型预测仅能解释3.05%的剩余个体差异,远低于人类重测信度53.6%。无论使用更丰富的个人资料、不同模型变体或微调,都无法缩小此差距。方差分析显示,移除项目均值后,可靠的剩余信号为‘人-项目’交互效应,其规模是稳定‘人’效应的8.9倍。个人资料编码了受访者特征,但未捕捉到这一项目特异性偏差。大模型还压缩了人类响应分布,使用更窄的范围、更少的类别和失真的分布形态。我们称此现象为‘项目均值代理’。当前的大模型代理可近似项目均值,但无法替代个体人类所需的分布特征与个性化偏差。我们提出四项实证测试以验证基于大模型的人类代理主张。

原文摘要 · Abstract (English)

LLMs are increasingly used as human surrogates, often on the premise that richer persona data could make them substitutes or exploratory tools for specific individuals. We test this premise across four datasets covering more than 400,000 participants and more than 6,000 survey items and experimental outcomes. LLMs perform well at the aggregate level: their average responses closely align with average human responses to the same items. But this success largely reflects predicting each item's average human response. Once each item's human mean is removed, LLM predictions explain only 3.05% of the remaining respondent-specific variation, far below the 53.6% human test-retest benchmark. Richer personas, model variants, and fine-tuning do not close this gap. In variance analyses, once item means are removed, the reliable remaining signal is person-by-item. It captures how a respondent departs from the mean on a particular item and is about 8.9x larger than the stable person effect. Persona data encode the respondent, but not this item-specific deviation. LLM responses also compress human response distributions, using less spread, fewer response categories, and distorted distributional shapes. We call this pattern item-mean surrogacy. Current LLM surrogates can approximate item averages, but not the distributions or respondent-specific deviations needed to replace individual humans. We propose four empirical tests for LLM-based human-surrogate claims.

大模型人类代理个体差异响应分布

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。