arXiv:2609.08637cs.CLcs.CY2026-09

测试大模型政治倾向,发现指令和语言显著影响结果

Navigating the digital spectrum: Assessing political bias, stability, and downstream fairness in Large Language Models

论文配图:Navigating the digital spectrum: Assessing political bias, stability, and downstream fairness in Large Language Models
图 1 · 摘自论文原文
  • 构建八维扰动框架,系统评估模型政治倾向
  • 多数模型平均偏自由左翼,但结果受指令等变量影响
  • 适合关注大模型公平性与评测可信度的研究者

大型语言模型日益作为信息中介被部署,但其政治行为的测量仍存在脆弱性,问卷结果常混杂模型倾向与测量偏差。本文提出稳健的「政治光谱测试」评估框架,通过在八个维度(语言、框架、指令、答案格式、选项顺序、人物设定表述)上采样300种配置,评估8个Gemma 3和Qwen 3模型在14种语言和3种量化级别下的表现,获得带不确定性的平均政治坐标。多数模型平均偏向自由左翼,但指令措辞、语言和答案格式显著影响坐标。跨语言差异主要源于坐标漂移而非文化推理差异。反向工程揭示轴权重失衡及退化响应向中心坍缩,小模型近原点估计可能反映弱信号而非真正中立。自由文本推理与聊天后分类会改变坐标,大模型表现出更清晰的人物区分,但特定的威权左翼人物设定未能有效引导多数模型向预期社会方向移动。下游任务中,人物设定影响较小,模型规模与目标群体对仇恨言论检测的影响更大;基础与中立提示在主题情感判断上一致性最高。因此,政治角色提示具有可测量但任务和数据集依赖的下游效应。

原文摘要 · Abstract (English)

Large Language Models are increasingly deployed as information intermediaries, yet measuring their political behavior remains fragile because questionnaire results mix model dispositions with measurement artifacts and response-elicitation biases. We introduce a robust Political Compass Test evaluation framework that samples 300 configurations across an eight-dimensional perturbation space varying language, framing, instructions, answer format, option order, and persona wording. We evaluate eight Gemma 3 and Qwen 3 models across 14 languages and three quantization levels, obtaining design-averaged political coordinates with quantified uncertainty. Most models lean Libertarian-Left on average, but instruction phrasing, language, and answer format significantly affect recovered coordinates. Cross-lingual differences primarily reflect coordinate drift rather than distinct cultural reasoning. Reverse-engineering the test also exposes axis-weighting imbalances and the collapse of degenerate responses toward the center, so near-origin estimates for the smallest models can reflect weak signal rather than centrism. Free-text reasoning and chat-then-classify elicitation alter recovered coordinates, and larger models show clearer persona separation, with a specific failure of the Authoritarian-Left persona to move most models in the intended social direction. In downstream tasks, persona effects are modest relative to model size and target group for hate-speech detection, while base and centrist prompts give the highest agreement for topic-level sentiment. Political role prompting therefore has measurable but task- and dataset-specific downstream effects.

大模型评测政治偏见公平性多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。