arXiv:2412.14843cs.CLcs.AI2024-12被引 19

用虚拟人格测试大模型政治倾向,发现其倾向左自由主义且易被右威权引导。

Mapping and Influencing the Political Ideology of Large Language Models using Synthetic Personas

  • 通过虚拟人格提示,用政治光谱测试映射模型倾向分布。
  • 多数模型集中在左自由主义区,右威权引导效果显著,左自由引导较弱。
  • 揭示模型对意识形态操控的不对称响应,适合研究偏见与可控生成的人看。

大型语言模型的政治偏见分析通常将其视为具有固定观点的单一实体。尽管已有多种测量方法,但基于人格提示对模型政治取向的影响仍未知。本文利用PersonaHub中的合成人格描述,通过政治光谱测试(PCT)映射人格提示下模型的政治分布,并检验是否可通过明确的意识形态提示,将其引导至截然相反的政治方向:右威权与左自由主义。实验显示,合成人格主导集中在左自由主义象限;当使用明确意识形态提示时,所有模型均显著转向右威权立场,但向左自由主义的转变受限,表明模型对意识形态操控存在不对称响应,可能反映训练数据中的固有偏见。

原文摘要 · Abstract (English)

The analysis of political biases in large language models (LLMs) has primarily examined these systems as single entities with fixed viewpoints. While various methods exist for measuring such biases, the impact of persona-based prompting on LLMs' political orientation remains unexplored. In this work we leverage PersonaHub, a collection of synthetic persona descriptions, to map the political distribution of persona-based prompted LLMs using the Political Compass Test (PCT). We then examine whether these initial compass distributions can be manipulated through explicit ideological prompting towards diametrically opposed political orientations: right-authoritarian and left-libertarian. Our experiments reveal that synthetic personas predominantly cluster in the left-libertarian quadrant, with models demonstrating varying degrees of responsiveness when prompted with explicit ideological descriptors. While all models demonstrate significant shifts towards right-authoritarian positions, they exhibit more limited shifts towards left-libertarian positions, suggesting an asymmetric response to ideological manipulation that may reflect inherent biases in model training.

大模型偏见人格提示政治倾向可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。