arXiv:2503.16148cs.CYcs.CL2025-03ACL被引 21

基于政治学理论,量化大模型的政治倾向,发现指令微调模型普遍偏左。

Only a Little to the Left: A Theory-grounded Measure of Political Bias in Large Language Models

  • 用政治学理论设计问卷,测试多种提示方式下的模型反应。
  • 分析88,110条响应,发现指令微调模型总体更偏左,但结果易受提示影响。
  • 开源代码数据,适合研究模型偏差与社会影响的学者使用。

GPT4、LLaMa等基于提示的语言模型被广泛用于模拟智能体、信息检索和内容分析等场景,其政治偏见可能影响性能。现有研究多采用政治光谱测试(PCT)评估偏见,但该工具缺乏科学有效性,且提示方法不一导致结果不一致,多数研究依赖固定答案设置。本文基于政治学理论,结合调查设计原则,构建可衡量政治偏见的新方法,测试多种提示输入并考虑提示敏感性。我们对11个开源与商用模型(含指令微调与非指令微调版本)进行测试,自动分类88,110条响应的政治立场。结果显示,尽管PCT会夸大某些模型如GPT3.5的偏见,整体偏见测量不稳定,但指令微调模型普遍呈现更明显的左倾倾向。相关代码与数据已公开于GitHub。

原文摘要 · Abstract (English)

Prompt-based language models like GPT4 and LLaMa have been used for a wide variety of use cases such as simulating agents, searching for information, or for content analysis. For all of these applications and others, political biases in these models can affect their performance. Several researchers have attempted to study political bias in language models using evaluation suites based on surveys, such as the Political Compass Test (PCT), often finding a particular leaning favored by these models. However, there is some variation in the exact prompting techniques, leading to diverging findings, and most research relies on constrained-answer settings to extract model responses. Moreover, the Political Compass Test is not a scientifically valid survey instrument. In this work, we contribute a political bias measured informed by political science theory, building on survey design principles to test a wide variety of input prompts, while taking into account prompt sensitivity. We then prompt 11 different open and commercial models, differentiating between instruction-tuned and non-instruction-tuned models, and automatically classify their political stances from 88,110 responses. Leveraging this dataset, we compute political bias profiles across different prompt variations and find that while PCT exaggerates bias in certain models like GPT3.5, measures of political bias are often unstable, but generally more left-leaning for instruction-tuned models. Code and data are available on: https://github.com/MaFa211/theory_grounded_pol_bias

模型偏见政治倾向提示工程大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。