LLM评价政策时会受支持国身份影响,西方背书更易获高分
Geopolitical alignment: Endorsement effects in large language models

- 用随机国家背书同一政策,测试LLM评分差异
- 中国/俄罗斯背书的政策得分显著低于欧美背书版本
- 要求解释理由后,部分模型对非西方背书惩罚加剧
大型语言模型(LLMs)越来越多地用于总结和评估具有政策相关性的信息,但其判断是否受到地缘政治线索的隐性影响尚不明确。本文通过一项背书实验,让四个LLM在政策被随机标注为美国、欧盟、中国或俄罗斯支持的情况下,评估相同的国际经济与安全政策。在仅输出评分的条件下,GPT-5、Claude Sonnet和Gemini对由中国或俄罗斯支持的政策评分明显低于由美国或欧盟支持的相同政策;DeepSeek是主要例外。第二阶段要求模型在评分后提供简短理由,该请求使GPT-5和Claude Sonnet的西方/非西方评分差距依然存在,弱化了Gemini的惩罚效应,并显著强化了DeepSeek对中、俄背书政策的负面评价。分析表明,西方背书常被视为可信度信号,而中、俄背书则被解读为数据安全、主权、监控或地缘政治风险的信号。研究揭示,即使政策内容完全一致,外国背书者身份仍可能影响LLM的政策评估结果。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used to summarize and evaluate policy-relevant information, but it remains unclear whether their judgments are implicitly shaped by geopolitical cues. I study this question with an endorsement experiment in which four LLMs evaluate the same international economic and security policies after each policy is randomly described as supported by the United States, the European Union, China, or Russia. In the numeric-only condition, GPT-5, Claude Sonnet, and Gemini rate China- and Russia-endorsed policies substantially lower than identical policies endorsed by the United States or the European Union; DeepSeek is the main exception. A second condition asks models to provide a short justification with the score. This request leaves the broad Western/non-Western gap intact for GPT-5 and Claude Sonnet, attenuates Gemini's penalties, and sharply activates China and Russia penalties in DeepSeek. The justifications indicate that Western endorsement is often treated as a credibility cue, whereas Chinese and Russian endorsement is treated as a cue for data security, sovereignty, surveillance, or geopolitical risk. These findings show that LLM policy evaluations can depend on the identity of a foreign endorser even when policy content is held fixed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。