用大模型模拟问卷回答,发现其难以真实反映不同人群的多元观点。
An Analysis of Large Language Models for Simulating User Responses in Surveys
- 提出CLAIMSIM方法,通过上下文引入多元观点增强响应多样性。
- 实验显示两种方法均无法准确模拟用户,平均准确率低于40%。
- 大模型固守单一立场,难区分群体差异,适合研究偏见而非真实用户。
使用大型语言模型(LLMs)模拟用户意见受到越来越多关注。然而,尤其是经过人类反馈强化学习(RLHF)训练的模型,常表现出对主流观点的偏倚,引发对其能否代表不同人口与文化背景用户的担忧。本文通过直接提示和思维链提示,评估了LLMs在跨领域问卷问答任务中模拟人类响应的能力。进一步提出名为CLAIMSIM的方法,从模型参数知识中提取多元观点作为上下文输入。实验表明,尽管CLAIMSIM生成了更多样化的回应,但两种方法在准确性上仍表现不佳。深入分析揭示两个关键局限:(1)模型在不同人口统计特征下保持固定立场,仅生成单一视角;(2)面对冲突观点时,模型难以推理人口特征间的细微差异,限制了对特定用户画像的响应适配能力。
原文摘要 · Abstract (English)
Using Large Language Models (LLMs) to simulate user opinions has received growing attention. Yet LLMs, especially trained with reinforcement learning from human feedback (RLHF), are known to exhibit biases toward dominant viewpoints, raising concerns about their ability to represent users from diverse demographic and cultural backgrounds. In this work, we examine the extent to which LLMs can simulate human responses to cross-domain survey questions through direct prompting and chain-of-thought prompting. We further propose a claim diversification method CLAIMSIM, which elicits viewpoints from LLM parametric knowledge as contextual input. Experiments on the survey question answering task indicate that, while CLAIMSIM produces more diverse responses, both approaches struggle to accurately simulate users. Further analysis reveals two key limitations: (1) LLMs tend to maintain fixed viewpoints across varying demographic features, and generate single-perspective claims; and (2) when presented with conflicting claims, LLMs struggle to reason over nuanced differences among demographic features, limiting their ability to adapt responses to specific user profiles.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。