arXiv:2502.16280cs.LGcs.AI2025-02被引 5

LLM生成的模拟民意数据在人群差异上失真,预测效果不可靠。

Human Preferences in Large Language Model Latent Space: A Technical Analysis on the Reliability of Synthetic Data in Voting Outcome Prediction

  • 通过探测模型隐空间,分析不同人口属性和提示对意见映射的影响。
  • 14个模型显示合成数据无法复现真实人群中的观点分布差异。
  • 提示敏感性高,导致结果不稳定,不适合用于社会科学研究。

生成式AI(GenAI)正被广泛用于模拟人类偏好以开展调查研究。尽管现有研究常通过比较模型生成回答与真实调查结果来评估合成数据质量,但关于大型语言模型(LLMs)能否作为人类受访者替代品的根本问题仍待解答。本研究对14种不同模型进行了技术分析,探讨人口属性与提示变化如何影响LLMs中的潜在意见映射,并评估其在基于调查的预测中的适用性。结果发现,LLM生成的数据无法再现真实世界中的人群观点方差,尤其是在不同人口子群体间。在政治领域,角色到政党的映射区分度有限,导致合成数据缺乏真实调查中观察到的观点精细分布。此外,我们发现某些模型对提示高度敏感,显著改变输出结果,进一步削弱了基于LLM模拟的稳定性与预测能力。作为核心贡献,我们采用基于探测的方法揭示了LLMs在其隐空间中编码政治归属的方式,暴露了这些模型引入的系统性偏差。研究警示:在公共舆论研究、社会科学实验及计算行为建模中使用人工智能生成的调查数据需保持谨慎。

原文摘要 · Abstract (English)

Generative AI (GenAI) is increasingly used in survey contexts to simulate human preferences. While many research endeavors evaluate the quality of synthetic GenAI data by comparing model-generated responses to gold-standard survey results, fundamental questions about the validity and reliability of using LLMs as substitutes for human respondents remain. Our study provides a technical analysis of how demographic attributes and prompt variations influence latent opinion mappings in large language models (LLMs) and evaluates their suitability for survey-based predictions. Using 14 different models, we find that LLM-generated data fails to replicate the variance observed in real-world human responses, particularly across demographic subgroups. In the political space, persona-to-party mappings exhibit limited differentiation, resulting in synthetic data that lacks the nuanced distribution of opinions found in survey data. Moreover, we show that prompt sensitivity can significantly alter outputs for some models, further undermining the stability and predictiveness of LLM-based simulations. As a key contribution, we adapt a probe-based methodology that reveals how LLMs encode political affiliations in their latent space, exposing the systematic distortions introduced by these models. Our findings highlight critical limitations in AI-generated survey data, urging caution in its use for public opinion research, social science experimentation, and computational behavioral modeling.

大模型民意模拟合成数据偏见探测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。