arXiv:2602.18462cs.CYcs.AI2026-02综述被引 6

用人物设定生成问卷回答,反而让部分群体数据失真。

Assessing the Reliability of Persona-Conditioned LLMs as Synthetic Survey Respondents

  • 用人物角色设定让大模型回答问卷,测试其一致性。
  • 超过7万次回答中,角色设定未提升整体准确性,反而加剧偏差。
  • 少数问题和小众群体受影响最严重,影响分析可靠性。

将人物设定驱动的大语言模型用作合成问卷受访者已成为计算社会科学与基于代理的模拟中的常见做法。然而,多属性人物设定是否真正提升模型可靠性仍不明确。本文利用来自世界价值观调查的美国微观数据集,评估了两种开源聊天模型与一个随机猜测基线,在超过7万例受试者-题目实例上的表现。结果发现,人物设定并未带来明显的整体对齐度提升,且在多数情况下显著降低性能。人物效应高度异质:大多数题目变化微小,但少数题目及代表性不足的子群体遭受不成比例的扭曲。研究揭示当前基于人物的模拟实践存在关键负面效应:人口统计条件化可能重新分布误差,损害子群体真实性,进而误导下游分析。

原文摘要 · Abstract (English)

Using persona-conditioned LLMs as synthetic survey respondents has become a common practice in computational social science and agent-based simulations. Yet, it remains unclear whether multi-attribute persona prompting improves LLM reliability or instead introduces distortions. Here we contribute to this assessment by leveraging a large dataset of U.S. microdata from the World Values Survey. Concretely, we evaluate two open-weight chat models and a random-guesser baseline across more than 70K respondent-item instances. We find that persona prompting does not yield a clear aggregate improvement in survey alignment and, in many cases, significantly degrades performance. Persona effects are highly heterogeneous as most items exhibit minimal change, while a small subset of questions and underrepresented subgroups experience disproportionate distortions. Our findings highlight a key adverse impact of current persona-based simulation practices: demographic conditioning can redistribute error in ways that undermine subgroup fidelity and risk misleading downstream analyses.

大模型评估合成数据问卷模拟偏差分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。