LLM在问卷调查中易受提示词干扰,表现出类似人类的偏好偏差。
Prompt Perturbations Reveal Human-Like Biases in Large Language Model Survey Responses
- 通过10类提示扰动测试9个大模型对问卷的响应
- 所有模型均显示显著末位选项偏好,且小模型更敏感
- 适合研究合成数据可靠性或提示工程的学者参考
大型语言模型(LLMs)正被越来越多地用作社会科学调查中人类受试者的替代,但其可靠性及对已知人类响应偏差(如中心化倾向、观点漂移、首因偏差)的敏感性仍不明确。本研究在世界价值观调查(WVS)的规范问卷语境下,测试了九个LLMs的响应鲁棒性,对问题表述和答案选项结构施加了十类综合扰动,共生成超过16.7万次模拟问卷访谈。结果不仅揭示了LLMs对扰动的脆弱性,还发现所有测试模型均存在稳定的末位偏好,即更倾向于选择最后呈现的答案选项。尽管更大模型通常更具鲁棒性,但所有模型仍对语义变化(如改写)及组合扰动敏感。这凸显了在使用LLMs生成合成调查数据时,提示设计与鲁棒性测试的重要性。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly used as proxies for human subjects in social science surveys, but their reliability and susceptibility to known human-like response biases, such as central tendency, opinion floating and primacy bias are poorly understood. This work investigates the response robustness of LLMs in normative survey contexts, we test nine LLMs on questions from the World Values Survey (WVS), applying a comprehensive set of ten perturbations to both question phrasing and answer option structure, resulting in over 167,000 simulated survey interviews. In doing so, we not only reveal LLMs' vulnerabilities to perturbations but also show that all tested models exhibit a consistent recency bias, disproportionately favoring the last-presented answer option. While larger models are generally more robust, all models remain sensitive to semantic variations like paraphrasing and to combined perturbations. This underscores the critical importance of prompt design and robustness testing when using LLMs to generate synthetic survey data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。