语言模型的地缘偏见源于后训练阶段,而非预训练数据。
It's the humans, not the data: Geopolitical bias in LLMs originates in post-training, amplified by the language of the prompt

- 通过对比基础模型与聊天模型,发现偏见在后训练中产生。
- 阿里Qwen 2.5对中国的偏好从-0.15升至+2.91,增幅达18倍。
- 提示语语言影响偏见方向,如法语提示强化法国倾向。
普遍认为语言模型的地缘偏见来自预训练数据。我们测试了来自七家实验室的七组开源大模型(基础模型与聊天模型),在英、法、中三语环境下对28个国家组合进行强制选择探测,发现地缘偏见实际产生于后训练阶段而非预训练。七家机构中六家在后训练后出现与开发者所属国家相关的偏好偏移。其中阿里Qwen 2.5的表现最为显著:基础模型对中国的偏好为-0.15对数优势(p=0.15),而聊天模型变为+2.91(p<10^-4),对数优势提升18倍。所有模型均出现对其他国家的偏好变化。此外,提示语语言显著影响偏见强度——法语开发的Mistral仅在法语提示下表现出亲法国倾向(FR-EN差值+1.91,p<10^-4)。研究揭示,地缘偏好并非源自互联网数据,而是由后训练过程主动塑造,凸显对对齐流程透明度、审计与监督的迫切需求。
原文摘要 · Abstract (English)
It has generally been assumed that geopolitical bias in language models originates from the training data used during the pre-training phase. We tested seven open-weight LLM pairs consisting of the base model (pre-training only) and the chat model (pre-training and post-training) from seven labs on a paired-scenario forced-choice probe over 28 country pairs in English, French, and Chinese, and found that geopolitical bias originates in post-training rather than in pre-training. Across seven AI labs, six showed shifts in the direction associated with the country or region of the model developer after post-training. This shift is strongest in Alibaba's Qwen 2.5: while the base is neutral on China-favourability (-0.15 log-odds, p=0.15), the post-trained chat variant is at +2.91 (p<10^-4), an 18x shift in odds. We also observe shifts in biases toward other countries across all models. Additionally, the magnitude of this shift depends on the language used to prompt the model: the French-made Mistral becomes pro-France only under French prompting (FR-EN shift +1.91, p<10^-4). These findings suggest that geopolitical preferences in language models are not simply inherited from large-scale internet data but are actively shaped during post-training, highlighting the need for greater transparency, auditing, and oversight of alignment processes that influence how models represent nations, cultures, and political perspectives.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。