用大模型生成心理量表模拟数据,可验证群体结构但难替代真实个体数据。
In Silico Development of Psychometric Scales: Feasibility of Representative Population Data Simulation with LLMs
- 让大模型扮演不同人群回答量表,生成虚拟被试数据
- 三组研究中模拟数据复现了真实量表的因子结构
- 适合早期量表设计阶段,不适用于个体层面验证
开发和验证心理量表需大量样本、多轮测试及资源投入。近期大语言模型(LLMs)的发展使通过提示模型以特定人口特征身份答题,生成合成被试数据成为可能,或可在真实数据收集前实现虚拟预演。本研究开展四项预注册研究(每项约 N=300),检验LLM生成的数据是否能复现人类响应的潜在结构与测量特性。研究1-2比较了模拟数据与真实数据在两个已验证量表上的表现;研究3-4基于模拟数据进行探索性因子分析构建新量表,并检验其在新收集的人类样本中的泛化能力。结果显示,在四次研究中有三次模拟数据成功再现了预期因子结构,且表现出一致的配置不变性与度量不变性,两个新量表还实现了标量不变性。然而,相关性检验显示真实与合成数据间存在显著差异,得分分布与方差亦有明显偏差。总体而言,尽管LLM能捕捉群体层面的潜在结构,但无法逼近个体层面的数据特征。模拟数据在性别上呈现完全内部不变性。因此,LLM生成数据适用于早期阶段的群体水平心理量表原型设计,但不能替代个体层面的验证。本文讨论了方法局限、偏倚风险与数据污染问题,以及相关伦理考量。
原文摘要 · Abstract (English)
Developing and validating psychometric scales requires large samples, multiple testing phases, and substantial resources. Recent advances in Large Language Models (LLMs) enable the generation of synthetic participant data by prompting models to answer items while impersonating individuals of specific demographic profiles, potentially allowing in silico piloting before real data collection. Across four preregistered studies (N = circa 300 each), we tested whether LLM-simulated datasets can reproduce the latent structures and measurement properties of human responses. In Studies 1-2, we compared LLM-generated data with real datasets for two validated scales; in Studies 3-4, we created new scales using EFA on simulated data and then examined whether these structures generalized to newly collected human samples. Simulated datasets replicated the intended factor structures in three of four studies and showed consistent configural and metric invariance, with scalar invariance achieved for the two newly developed scales. However, correlation-based tests revealed substantial differences between real and synthetic datasets, and notable discrepancies appeared in score distributions and variances. Thus, while LLMs capture group-level latent structures, they do not approximate individual-level data properties. Simulated datasets also showed full internal invariance across gender. Overall, LLM-generated data appear useful for early-stage, group-level psychometric prototyping, but not as substitutes for individual-level validation. We discuss methodological limitations, risks of bias and data pollution, and ethical considerations related to in silico psychometric simulations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。