arXiv:2511.03699cs.CLcs.CY2025-11被引 2

测试大模型是否会有阴谋思维,发现它们易被诱导且存在偏见。

Do Androids Dream of Unseen Puppeteers? Probing for a Conspiracy Mindset in Large Language Models

  • 用心理量表测试大模型的阴谋心态,对比不同提示效果。
  • 模型部分认同阴谋论,社会属性提示会引发不均衡偏差。
  • 针对性提示可轻易引导模型走向阴谋论,适合安全与伦理研究者关注。

本文研究大语言模型(LLMs)是否具有阴谋倾向,是否存在社会人口学偏差,以及其在何种条件下易被引导形成阴谋论观点。阴谋信念是虚假信息传播和对制度不信任的核心因素,因此是评估大模型社会拟合度的关键场景。尽管大模型常被用作人类行为的代理,但对其是否复现高级心理建构如阴谋心态仍知之甚少。为此,我们对多个模型施以经验证的心理量表,考察不同提示与条件化策略下的表现。结果表明,大模型在一定程度上与阴谋信念元素一致,且社会人口学特征的条件化导致非均匀影响,暴露潜在的群体偏差。此外,特定提示可轻易使模型回应转向阴谋方向,凸显其易受操控性及在敏感场景中部署的风险。这些发现强调了需从心理学维度批判性评估大模型,既推动计算社会科学的发展,也为应对有害应用提供缓解策略。

原文摘要 · Abstract (English)

In this paper, we investigate whether Large Language Models (LLMs) exhibit conspiratorial tendencies, whether they display sociodemographic biases in this domain, and how easily they can be conditioned into adopting conspiratorial perspectives. Conspiracy beliefs play a central role in the spread of misinformation and in shaping distrust toward institutions, making them a critical testbed for evaluating the social fidelity of LLMs. LLMs are increasingly used as proxies for studying human behavior, yet little is known about whether they reproduce higher-order psychological constructs such as a conspiratorial mindset. To bridge this research gap, we administer validated psychometric surveys measuring conspiracy mindset to multiple models under different prompting and conditioning strategies. Our findings reveal that LLMs show partial agreement with elements of conspiracy belief, and conditioning with socio-demographic attributes produces uneven effects, exposing latent demographic biases. Moreover, targeted prompts can easily shift model responses toward conspiratorial directions, underscoring both the susceptibility of LLMs to manipulation and the potential risks of their deployment in sensitive contexts. These results highlight the importance of critically evaluating the psychological dimensions embedded in LLMs, both to advance computational social science and to inform possible mitigation strategies against harmful uses.

大模型安全阴谋心理提示工程社会偏见

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。