用大模型生成问卷答案,发现限制输出法最准
Survey Response Generation: Generating Closed-Ended Survey Responses In-Silico with Large Language Models
- 对比8种生成方法,限制输出效果最佳
- 3200万次模拟显示不同方法差异显著
- 适合做社会调查仿真的研究者参考
大量基于大语言模型(LLMs)的问卷响应仿真研究聚焦于生成封闭式问卷回答,而传统LLM训练目标为开放文本生成。此前研究采用多种方法生成封闭式响应,尚未形成标准实践。本文系统评估了8种问卷响应生成方法在4份政治态度问卷和10个开源大模型上的表现,共生成3200万条模拟响应。结果表明,个体层面与子群体层面的响应对齐度存在显著差异;受限生成方法整体表现最优,而推理输出并未持续提升对齐度。研究强调生成方法对仿真结果的重大影响,提出实际应用建议。
原文摘要 · Abstract (English)
Many in-silico simulations of human survey responses with large language models (LLMs) focus on generating closed-ended survey responses, whereas LLMs are typically trained to generate open-ended text instead. Previous research has used a diverse range of methods for generating closed-ended survey responses with LLMs, and a standard practice remains to be identified. In this paper, we systematically investigate the impact that various Survey Response Generation Methods have on predicted survey responses. We present the results of 32 mio. simulated survey responses across 8 Survey Response Generation Methods, 4 political attitude surveys, and 10 open-weight language models. We find significant differences between the Survey Response Generation Methods in both individual-level and subpopulation-level alignment. Our results show that Restricted Generation Methods perform best overall, and that reasoning output does not consistently improve alignment. Our work underlines the significant impact that Survey Response Generation Methods have on simulated survey responses, and we develop practical recommendations on the application of Survey Response Generation Methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。