用大模型生成模拟数据时,分析选择不同会导致结果差异巨大。
The threat of analytic flexibility in using large language models to simulate human data
- 通过调整模型、提示词等配置生成252种模拟样本
- 不同配置下与真实数据的相关性从0.23到0.84不等
- 提醒研究者警惕分析灵活性对模拟数据可信度的影响
社会科学家正使用大语言模型生成“硅样本”:旨在替代人类受访者的合成数据集。但生成这些样本需做出诸多分析决策,包括模型选择、采样参数、提示格式以及提供的人口或情境信息量。在两项研究中,我检验了这些选择是否显著影响硅样本与人类数据的一致性。研究1中,针对两个社会心理学量表,生成252种硅样本配置,评估其是否恢复了参与者排名、响应分布及量表间相关性。结果显示,各配置在三个维度上差异显著,某些配置在一个指标上表现良好,但在其他指标上表现较差。研究2将分析扩展至一篇已发表的硅样本使用案例,重新审视Argyle等(2023)研究3,采用66种替代配置。人类与硅样本关联结构间的相关性在不同配置间变化极大,从r = .23到r = .84。综合来看,这些结果表明,尽管配置选择合理,仍可能显著改变对硅样本保真度的结论。我呼吁关注分析灵活性带来的威胁,并提出研究人员可采取的缓解策略。
原文摘要 · Abstract (English)
Social scientists are now using large language models to create "silicon samples": synthetic datasets intended to stand in for human respondents. However, producing these samples requires many analytic choices, including model selection, sampling parameters, prompt format, and the amount of demographic or contextual information provided. Across two studies, I examine whether these choices materially affect correspondence between silicon samples and human data. In Study 1, I generated 252 silicon-sample configurations for a controlled case study using two social-psychological scales, evaluating whether configurations recovered participant rankings, response distributions, and between-scale correlations. Configurations varied substantially across all three criteria, and configurations that performed well on one dimension often performed poorly on another. In Study 2, I extended this analysis to a published silicon-sample use case by re-examining Argyle et al.'s (2023) Study 3 using 66 alternative configurations. Correlations between human and silicon association structures differed substantially across configurations, from r = .23 to r = .84. Taken together, the results from these studies demonstrate that different defensible configuration choices can materially alter conclusions about the fidelity of silicon samples. I call for greater attention to the threat of analytic flexibility in using silicon samples and outline strategies that researchers may adopt to reduce this threat.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。