arXiv:2511.10695cs.CL2025-11中稿 · AAAI被引 1

揭示大模型在国际关系中的国家偏见,提出有效减偏方法。

"As Eastern Powers, I will veto." : An Investigation of Nation-level Bias of Large Language Models in International Relations

  • 通过联合国安理会历史数据设计三类测试框架
  • 不同模型对五大常任理事国偏见程度差异显著,且随任务变化
  • 结合检索与自我反思的去偏框架提升GPT-4o-mini和Llama-3.3-70B表现

本文系统研究大型语言模型(LLMs)在国际关系(IR)领域中存在的国家层面偏见。基于联合国安全理事会(UNSC)的历史记录,我们构建了一个包含三项测试的偏见评估框架,重点考察五常成员国的偏见表现。实验表明,尽管多数模型存在对西方国家的偏好及对俄罗斯的负面倾向,但具体偏差模式在不同模型间存在差异;更值得注意的是,同一模型对同一国家的偏见方向与强度会随评估上下文改变,表明模型偏见具有多维性,依赖于模型与任务。此外,推理能力越强的模型表现出更低的偏见与更高性能。基于此发现,我们提出一种结合检索增强生成与反思式自省的去偏框架,实验显示其能有效降低国家偏见,并提升在GPT-4o-mini和Llama-3.3-70B上的表现。研究强调,在应用大模型处理国际关系问题时,必须同时评估其偏见与性能。

原文摘要 · Abstract (English)

This paper systematically examines nation-level biases exhibited by Large Language Models (LLMs) within the domain of International Relations (IR). Leveraging historical records from the United Nations Security Council (UNSC), we developed a bias evaluation framework comprising three distinct tests to explore nation-level bias in various LLMs, with a particular focus on the five permanent members of the UNSC. Experimental results show that, even with the general bias patterns across models (e.g., favorable biases toward the western nations, and unfavorable biases toward Russia), these still vary based on the LLM. Notably, even within the same LLM, the direction and magnitude of bias for a nation change depending on the evaluation context. This observation suggests that LLM biases are fundamentally multidimensional, varying across models and tasks. We also observe that models with stronger reasoning abilities show reduced bias and better performance. Building on this finding, we introduce a debiasing framework that improves LLMs' factual reasoning combining Retrieval-Augmented Generation with Reflexion-based self-reflection techniques. Experiments show it effectively reduces nation-level bias, and improves performance, particularly in GPT-4o-mini and LLama-3.3-70B. Our findings emphasize the need to assess nation-level bias alongside performance when applying LLMs in the IR domain.

大模型偏见国际关系去偏方法多维偏见

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。