用跨问卷迁移评估大模型模拟受访者,发现其预测准确率接近真实模型。
Silicon Sampling via Cross-Survey Transfer
- 让大模型根据已答问题预测另一组新问题答案,检验个体级预测能力。
- 零样本模型对未见问题预测准确率达52%,仅比监督模型低6个百分点。
- 发现政治态度比主权议题更易预测,且模型偏差与安全对齐影响复杂。
硅样本——利用大语言模型(LLMs)模拟人类调查回答者——已成为增强传统调查研究的有前景方法。然而,多数评估依赖分布比较而非个体级预测,可能混淆模式匹配与真实预测。我们提出跨问卷迁移评估框架:给定一个受访者在一组问题上的回答,要求模型预测其在同一次调查中完全不同问题上的回答。基于台湾选举与民主化研究(TEDS)2024数据,使用三个开源大模型(27B-120B参数)及监督学习基线,我们发现:(1)零样本大模型在真正未见问题上达到52%准确率,仅比同人群训练的监督随机森林低6个百分点;(2)存在稳定的可预测性层级,从政党认同的67%到主权议题的23%;(3)方差坍缩与安全对齐效应——常被视作大模型局限——比以往认知更复杂,方差坍缩同样影响监督模型,而对齐效果在不同模型家族间差异显著。这些发现厘清了硅样本的潜力与边界。
原文摘要 · Abstract (English)
Silicon sampling-using large language models (LLMs) to simulate human survey respondents-has emerged as a promising approach for augmenting traditional survey research. However, most evaluations rely on distributional comparisons rather than individual-level prediction, which risks conflating pattern matching with coherent respondent-level prediction. We propose cross-survey transfer, a more rigorous evaluation framework in which an LLM is given a respondent's answers to one set of questions and must predict their answers to entirely different questions from the same survey. Using data from the Taiwan Election and Democratization Study (TEDS) 2024, three open-weight LLMs (27B-120B parameters), and supervised machine learning baselines, we find that: (1) zero-shot LLMs achieve 52% accuracy on genuinely unseen items, closing to within 6 percentage points (pp) of a supervised random forest trained on same-population data; (2) a stable construct predictability hierarchy emerges, from 67% for partisan attitudes to 23% for sovereignty; and (3) variance collapse and safety alignment effects-two commonly cited LLM limitations-turn out to be more nuanced than previously reported, with variance collapse affecting supervised models as well and alignment effects varying dramatically across model families. These findings clarify both the promise and boundaries of silicon sampling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。