arXiv:2608.06123cs.AIcs.CL2026-08

通过对比各国对等冲突场景的回应差异,揭示大模型在国际政治中的偏见。

Poli-Bias: Understanding and Measuring Large Language Model Biases in International Political Conflicts

论文配图:Poli-Bias: Understanding and Measuring Large Language Model Biases in International Political Conflicts
图 1 · 摘自论文原文
  • 设计反事实框架,系统交换国家身份对比响应差异。
  • 13个主流大模型中均发现国家身份影响法律描述与评价。
  • 可细粒度检测模型在国际事务中的不公对待或谄媚倾向。

衡量大语言模型(LLMs)在国际政治冲突中的政治偏见仍具挑战性,因其可能通过微妙的表述、论证和法律推理差异体现,难以用单一指标捕捉。本文提出Poli-Bias,一种反事实框架,用于评估LLMs是否因涉及国家不同而对法律上等效的冲突情境作出不同处理。该框架通过系统交换多样地缘关系、法律违规行为和推理任务中的国家身份,比较成对提示下的响应。不同于将偏见简化为单一判断,Poli-Bias将响应差异分解为五个可解释维度,揭示不平等对待的具体表现。在涵盖多种模型家族和规模的13个当代大模型中,我们发现国家身份与用户归属会系统性影响等效行为在国际法下的描述、评价与辩护方式。结果表明,Poli-Bias可作为细粒度审计模型政治中立性与谄媚倾向的有效工具。

原文摘要 · Abstract (English)

Measuring political bias in large language models (LLMs) remains challenging as it can manifest through subtle differences in framing, argumentation, and legal reasoning that are difficult to capture with a single metric. In this work, we introduce Poli-Bias, a counterfactual framework for measuring whether LLMs treat legally equivalent conflict scenarios differently depending on the countries involved. Poli-Bias compares responses to paired prompts in which country identities are systematically swapped across diverse geopolitical relationships, legal violations, and reasoning tasks. Rather than reducing bias to a single judgment, our framework decomposes response disparities into five interpretable dimensions, revealing how and where unequal treatment manifests. Across 13 contemporary LLMs spanning diverse model families and sizes, we find that country identities and user affiliations can systematically affect how equivalent actions are described, evaluated, and defended under international law. Our results thus establish Poli-Bias as a fine-grained framework for auditing political even-handedness and sycophancy in LLMs.

大模型偏见国际政治反事实分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。