arXiv:2504.17052cs.CL2025-04被引 1

提出评估大模型政治立场稳定性的自动化框架,发现模型在不同话题下立场可能不一致。

PReSS: An Automated Black-Box Framework for Evaluating Political Stance Stability in LLMs

  • 通过联合分析模型与话题上下文,分类四类立场响应模式。
  • 9个主流模型在19个议题上表现差异大,部分模型左倾却在特定议题上稳定右倾。
  • 适用于需要控制意识形态输出的场景,如模型对齐与去偏干预。

现有大语言模型(LLMs)的政治偏见评估多将其输出简单归类为左倾或右倾。本文拓展这一视角,考察模型在不同议题上的意识形态倾向变化及其一致性,即立场稳定性。为此提出PReSS(Political Response Stability under Stress)——一种自动化黑盒评估框架,结合模型与话题上下文,将响应分为四类:稳定左倾、不稳定左倾、稳定右倾、不稳定右倾。在9个广泛使用的LLM上对19个政治议题进行测试,发现立场稳定性存在显著差异;例如,整体左倾的模型在某些议题上呈现稳定右倾行为。这凸显了需采用主题感知和细粒度的评估方式。此外,稳定性对可控生成与模型对齐具有实际意义:当模型被提示或微调以转换意识形态时,不稳定议题的立场更易改变,而稳定议题则抗拒修改。因此,将稳定性作为调节因素,为理解、评估和引导政治敏感模型行为提供了理论基础。

原文摘要 · Abstract (English)

Existing evaluations of political bias in large language models (LLMs) typically classify outputs as left- or right-leaning. We extend this perspective by examining how ideological tendencies vary across topics and how consistently models maintain their positions, a property we refer to as stability. To capture this dimension, we propose PReSS (Political Response Stability under Stress), an automated black-box framework that evaluates LLMs by jointly considering model and topic context, categorizing responses into four stance types: stable-left, unstable-left, stable-right, and unstable-right. Applying PReSS to 9 widely used LLMs across 19 political topics reveals substantial variation in stance stability; for instance, a model that is left-leaning overall can exhibit stable-right behavior on certain topics. This highlights the importance of topic-aware and fine-grained evaluation of political ideologies of LLMs. Moreover, stability has practical implications for controlled generation and model alignment: interventions such as debiasing or ideology reversal should explicitly account for stance stability. Our empirical analyses reveal that when models are prompted or fine-tuned to adopt the opposite ideology, unstable topic stances are more likely to change, whereas stable ones resist modification. Thus, treating stability as a moderating factor provides a principled foundation for understanding, evaluating, and guiding interventions in politically sensitive model behavior.

大模型评估政治偏见立场稳定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。