arXiv:2507.02919cs.CLcs.CY2025-07被引 4

大模型生成意见存在偏差,无法真实反映人群多样性。

ChatGPT is not A Man but Das Man: Representativeness and Structural Consistency of Silicon Samples Generated by Large Language Models

  • 用真实问卷数据对比大模型回答,发现结构不一致
  • 模型对少数群体观点严重低估,呈现观点同质化
  • 适合关注AI社会影响与政策推演的研究者

大型语言模型(如ChatGPT、Llama 3.1系列)被视作模拟人类意见的“硅样本”,但本研究指出其可能扭曲总体意见。通过对比美国2020年全国选举研究(ANES)中关于堕胎与非法移民的问题,发现大模型在不同人口分组层面存在响应准确率不一致(结构性不一致),且显著低估少数派观点(同质化)。研究提出“准确性优化假说”:模型倾向选择主流回答,导致观点趋同。这质疑了将聊天机器人作为人类调查数据直接替代品的合理性,可能强化偏见并误导政策制定。

原文摘要 · Abstract (English)

Large language models (LLMs) in the form of chatbots like ChatGPT and Llama are increasingly proposed as "silicon samples" for simulating human opinions. This study examines this notion, arguing that LLMs may misrepresent population-level opinions. We identify two fundamental challenges: a failure in structural consistency, where response accuracy doesn't hold across demographic aggregation levels, and homogenization, an underrepresentation of minority opinions. To investigate these, we prompted ChatGPT (GPT-4) and Meta's Llama 3.1 series (8B, 70B, 405B) with questions on abortion and unauthorized immigration from the American National Election Studies (ANES) 2020. Our findings reveal significant structural inconsistencies and severe homogenization in LLM responses compared to human data. We propose an "accuracy-optimization hypothesis," suggesting homogenization stems from prioritizing modal responses. These issues challenge the validity of using LLMs, especially chatbots AI, as direct substitutes for human survey data, potentially reinforcing stereotypes and misinforming policy.

大模型偏差社会模拟观点同质化人机对比

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。