大模型政治倾向可测可调,多语言实验证明越大越左
Multilingual Political Views of Large Language Models: Identification and Steering
- 用14种语言的11种改写句测试7个开源大模型的政治立场
- 模型越大越倾向自由左派,不同语言和模型家族差异明显
- 简单干预就能改变立场,适合政策制定者和安全研究者
大型语言模型(LLMs)在日常应用中日益普及,引发对其潜在政治影响的担忧。现有研究显示这些模型常表现出可测量的政治偏见,通常偏向自由或进步立场,但关键空白仍存:多数研究仅覆盖少数模型和语言,难以判断偏见是否普遍;且极少探讨能否主动控制。本文通过大规模研究现代开源指令微调大模型的政治倾向,评估了包括LLaMA-3.1、Qwen-3和Aya-Expanse在内的7个模型,覆盖14种语言,使用包含11个语义等价改写句的政治理解测试以确保测量稳健性。结果表明,模型规模越大,越趋向自由左翼立场,且在不同语言和模型族间存在显著差异。为进一步测试立场可操控性,我们采用一种简单的质心激活干预技术,成功在多语言环境下可靠地将模型回应导向其他意识形态立场。代码已公开于https://github.com/d-gurgurov/Political-Ideologies-LLMs。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used in everyday tools and applications, raising concerns about their potential influence on political views. While prior research has shown that LLMs often exhibit measurable political biases--frequently skewing toward liberal or progressive positions--key gaps remain. Most existing studies evaluate only a narrow set of models and languages, leaving open questions about the generalizability of political biases across architectures, scales, and multilingual settings. Moreover, few works examine whether these biases can be actively controlled. In this work, we address these gaps through a large-scale study of political orientation in modern open-source instruction-tuned LLMs. We evaluate seven models, including LLaMA-3.1, Qwen-3, and Aya-Expanse, across 14 languages using the Political Compass Test with 11 semantically equivalent paraphrases per statement to ensure robust measurement. Our results reveal that larger models consistently shift toward libertarian-left positions, with significant variations across languages and model families. To test the manipulability of political stances, we utilize a simple center-of-mass activation intervention technique and show that it reliably steers model responses toward alternative ideological positions across multiple languages. Our code is publicly available at https://github.com/d-gurgurov/Political-Ideologies-LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。