arXiv:2506.14997cs.CYcs.CL2025-06被引 1

用假设检验量化大模型与人类在选择题中的行为差异

Hypothesis Testing for Quantifying LLM-Human Misalignment in Multiple Choice Settings

  • 基于假设检验构建评估框架,判断大模型能否真实模拟人类选择
  • 发现主流大模型在敏感议题上无法准确模拟不同人群行为
  • 为社会科学研究中使用大模型提供严谨评估方法,适合研究者参考

随着大型语言模型(LLMs)在社会科学(如经济学、营销学)研究中的广泛应用,评估其是否能有效模拟人类行为变得至关重要。本文提出一种基于假设检验的量化框架,用于评估大模型在多项选择问卷设置下与真实人类行为之间的偏差。该框架可系统判断特定语言模型是否能有效模拟人类意见、决策及通过多项选择表达的一般行为。我们将其应用于一个流行的语言模型,测试其在多个公共调查中对不同人群(如不同种族、年龄、收入群体)意见的模拟能力,结果表明该模型在涉及争议性问题时,对所测试子群体的模拟效果不佳。这揭示了该语言模型与目标人群之间存在显著偏差,提醒研究者在社会科学研究中应避免简单使用大模型模拟人类,需建立更严谨的评估流程。

原文摘要 · Abstract (English)

As Large Language Models (LLMs) increasingly appear in social science research (e.g., economics and marketing), it becomes crucial to assess how well these models replicate human behavior. In this work, using hypothesis testing, we present a quantitative framework to assess the misalignment between LLM-simulated and actual human behaviors in multiple-choice survey settings. This framework allows us to determine in a principled way whether a specific language model can effectively simulate human opinions, decision-making, and general behaviors represented through multiple-choice options. We applied this framework to a popular language model for simulating people's opinions in various public surveys and found that this model is ill-suited for simulating the tested sub-populations (e.g., across different races, ages, and incomes) for contentious questions. This raises questions about the alignment of this language model with the tested populations, highlighting the need for new practices in using LLMs for social science studies beyond naive simulations of human subjects.

大模型评估社会科学研究行为模拟假设检验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。