研究发现大模型意识形态可被角色设定显著影响,且越大模型越易受引导。
Political Ideology Shifts in Large Language Models
- 用角色设定测试7个模型的意识形态倾向变化
- 70B以上模型隐性意识形态覆盖达49%,远超7-8B模型
- 右翼提示效果明显,左翼提示则导致小模型反应不一
大型语言模型(LLMs)在政治敏感场景中的应用日益广泛,引发对其意识形态偏见的担忧。本文通过政治光谱测试(62条陈述)作为标准化行为探针,考察七种开源指令微调模型(参数量7B-72B)在合成角色设定下的意识形态表达变化。三组实验共生成20万个人格化角色,产生超过2.6亿次模型响应。结果表明:(i) 更大模型具有更广的隐性意识形态覆盖范围,从7-8B模型的14-35%提升至70B+模型的最高49%;(ii) 显式意识形态提示引发显著且统计上显著的偏移,右翼威权提示使所有模型朝预期方向移动,且多数轴向效应更大;(iii) 左翼自由主义提示导致更异质化响应,其中四个7-8B模型中有三个出现反向经济立场,而所有70B+模型均按预期移动;(iv) 角色描述中的主题相关语义内容与意识形态空间中系统性、可解释的偏移相关。本研究揭示了角色设定影响模型输出的上游机制,但未检验此类偏移是否影响用户信念、决策或政治行为。结果表明,在英语提示和角色条件下的语言模型中存在生成层的意识形态可塑性,需在评估政治中立性、公平性和安全性时考虑交互因素。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly deployed in politically sensitive contexts, raising concerns about their susceptibility to ideological biases. In this work, we examine how synthetic persona conditioning shapes ideological expression across seven open-weight instruction-tuned models (7B-72B parameters) using the Political Compass Test (62 statements) as a standardized behavioral probe. Across three studies involving 200,000 synthetic personas and more than 260 million model responses, we analyze implicit and explicit malleability, as well as theme-associated variations. We find that: (i) larger models exhibit broader implicit ideological coverage, increasing from 14-35% for 7-8B models to up to 49% for 70B+ models; (ii) explicit ideological priming induces large and statistically significant shifts, with right-authoritarian cues moving all models in the intended direction and producing larger effects in most model-axis comparisons; (iii) left-libertarian priming produces more heterogeneous responses, including counter-directional economic shifts in three of four 7-8B models, while all 70B+ models move in the intended direction; and (iv) theme-associated semantic content in persona descriptions is linked to systematic and interpretable directional shifts in ideological space. While our results identify an upstream mechanism through which persona conditioning can alter model responses under a standardized ideological probe, we do not test whether such shifts affect users beliefs, decisions, or political behavior. Our findings are best understood as evidence of ideological malleability at the generation layer, highlighting the need to account for interactional factors when evaluating political neutrality, fairness, and safety in English-prompted, persona-conditioned language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。