分析大模型对美英苏中历史事件的偏见,发现其倾向特定国家叙事。
Geopolitical biases in LLMs: what are the "good" and the "bad" countries according to contemporary language models
- 构建中立事件描述与多国视角对比数据集
- 模型明显偏好特定国家叙事,简单提示去偏效果有限
- 揭示模型对身份标签敏感,适合研究偏见与伦理的学者
本文通过分析大语言模型对美国、英国、苏联和中国在具有冲突性国家视角的历史事件中的解读,评估其地缘政治偏见。我们引入一个包含中立事件描述和多国不同观点的新数据集。研究发现,模型存在显著的地缘政治偏见,倾向于支持特定国家的叙事。此外,简单的去偏提示在减少这些偏见方面效果有限。通过操纵参与者标签的实验显示,模型对归属关系敏感,有时会放大偏见或识别出不一致之处。本工作揭示了大模型中的国家叙事偏见,挑战了简单去偏方法的有效性,并为未来地缘政治偏见研究提供了框架和数据集。
原文摘要 · Abstract (English)
This paper evaluates geopolitical biases in LLMs with respect to various countries though an analysis of their interpretation of historical events with conflicting national perspectives (USA, UK, USSR, and China). We introduce a novel dataset with neutral event descriptions and contrasting viewpoints from different countries. Our findings show significant geopolitical biases, with models favoring specific national narratives. Additionally, simple debiasing prompts had a limited effect in reducing these biases. Experiments with manipulated participant labels reveal models' sensitivity to attribution, sometimes amplifying biases or recognizing inconsistencies, especially with swapped labels. This work highlights national narrative biases in LLMs, challenges the effectiveness of simple debiasing methods, and offers a framework and dataset for future geopolitical bias research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。