arXiv:2608.14606cs.CYcs.AI2026-08综述

检验大模型生成的问卷数据是否真实反映人类心理特征

Plausible but Not Valid: A Psychometric Audit of LLMs as Synthetic Survey Respondents

论文配图:Plausible but Not Valid: A Psychometric Audit of LLMs as Synthetic Survey Respondents
图 1 · 摘自论文原文
  • 用心理测量学方法评估大模型问卷回答的真实性
  • 37个模型中无一超越统计基线,多数不如人类相似
  • 模型易产生虚假因果关系,不适合替代真实调查数据

大语言模型被越来越多地用作合成问卷响应者,但现有评估仅关注单个回答是否合理。本文提出应从心理测量学角度检验:大模型能否保留真实人类数据中的联合分布、潜在结构、信度、中介路径和人口统计效应。研究构建了立陶宛组织心理学数据集(n=263名员工;包含Dunham变革态度量表、UWES-17、Koopmans IWPQ,共68题,12个子量表),对37个模型(涵盖OpenAI、Anthropic、Google及12个开源模型家族)在五级人格披露梯度下,进行呈现方式、推理努力、反事实人口替换(性别、职位、教育)、跨语言验证和逐字记忆探测等测试。引入心理测量相似性评分(PSS),与5个非大模型统计基线及保留的人类-人类上限对比,使用响应者自助法置信区间和项目置换零模型检验Tucker's phi。结果显示,大模型虽能复现人类心理关系的定性方向,但高斯-柯普拉基线优于所有大模型;大模型群体内部相似性(均值0.73)高于与人类的相似性;记忆能力不主导排名(回忆-PSS相关性0.00)。反事实替换显示教育影响显著(|d|=0.56),远超性别(0.12)和职位(0.18);10个模型中有8个在UWES上Tucker's phi落在置换零模型内。下游分析表明,所有模型均存在强烈顺从偏移(+0.84标准差),合成训练回归器在保留人类样本上的预测有效性下降(平均R² -0.18 vs 0.28),且在10条安慰剂中介路径中有3条虚构出间接效应。大模型样本无法作为人类调查数据的直接替代品。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used as synthetic survey respondents, but existing evaluations ask whether answers look plausible at the individual level. We argue the right question is psychometric: do LLMs preserve the joint distribution, latent structure, reliability, mediation pathways, and demographic effects of real human survey data? We introduce a Lithuanian organisational-psychology dataset (n=263 employees; Dunham Attitudes Toward Change, UWES-17, Koopmans IWPQ; 68 items, 12 subscales) and condition a 37-model lineup spanning OpenAI, Anthropic, Google, and twelve open-weight families on real respondent profiles under a five-level persona-disclosure ladder, presentation and reasoning-effort ablations, counterfactual demographic swaps (gender, role, education), a cross-language check, and a verbatim-recall memorization probe. The resulting Psychometric Similarity Score (PSS) is anchored against five non-LLM statistical baselines and a held-out human-vs-human ceiling, with respondent-bootstrap confidence intervals and an item-permutation null for Tucker's phi. LLMs reproduce the qualitative direction of human psychometric relationships, but a Gaussian-copula baseline beats every LLM on the sample-driven PSS components; the LLM "crowd" is more similar to itself (mean inter-LLM PSS 0.73) than to humans; and memorization does not drive the leaderboard (recall-PSS rank correlation 0.00). Counterfactual swaps reveal education-driven effects (mean |d|=0.56) that dwarf gender (0.12) and role (0.18); Tucker's phi on UWES falls inside the permutation null for 8 of 37 models. Downstream, every LLM shows a strong acquiescence shift (+0.84 SD), synthetic-trained regressors lose predictive validity on held-out humans (mean R^2 -0.18 vs 0.28), and models fabricate indirect effects on 3 of 10 placebo mediation paths. LLM samples are not a drop-in replacement for human survey data.

心理测量大模型评估问卷生成合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。