揭示大模型在事实与敏感问题中的语言偏见,提出双层评估框架。
A Dual-Layered Evaluation of Geopolitical and Cultural Bias in LLMs
- 分两阶段评估:事实题一致性与地缘敏感议题响应。
- 查询语言影响事实回答,训练数据与语言共现敏感议题偏见。
- 适合关注多语言公平性与文化敏感性的研究者参考。
随着大语言模型(LLMs)在多种语言和文化背景中广泛应用,理解其在事实性与争议性场景下的表现至关重要,尤其当其输出可能影响公众意见或强化主流叙事时。本文定义两类偏见:模型偏见(源于训练数据)与推理偏见(由查询语言引发),并通过双阶段评估展开研究。第一阶段在存在唯一可验证答案的事实性问题上评估模型跨语言一致性;第二阶段扩展至地缘政治敏感争议问题,考察回答是否反映文化嵌入或意识形态倾向。我们构建了一个人工标注的数据集,涵盖四种语言及两类问题类型。结果显示,第一阶段呈现查询语言诱导的对齐现象,第二阶段则体现模型训练语境与查询语言的相互作用。本文提供了一种结构化框架,用于评估大模型在中立与敏感话题上的行为,为未来多语言部署及文化敏感性评估实践提供了洞见。
原文摘要 · Abstract (English)
As large language models (LLMs) are increasingly deployed across diverse linguistic and cultural contexts, understanding their behavior in both factual and disputable scenarios is essential, especially when their outputs may shape public opinion or reinforce dominant narratives. In this paper, we define two types of bias in LLMs: model bias (bias stemming from model training) and inference bias (bias induced by the language of the query), through a two-phase evaluation. Phase 1 evaluates LLMs on factual questions where a single verifiable answer exists, assessing whether models maintain consistency across different query languages. Phase 2 expands the scope by probing geopolitically sensitive disputes, where responses may reflect culturally embedded or ideologically aligned perspectives. We construct a manually curated dataset spanning both factual and disputable QA, across four languages and question types. The results show that Phase 1 exhibits query language induced alignment, while Phase 2 reflects an interplay between the model's training context and query language. This paper offers a structured framework for evaluating LLM behavior across neutral and sensitive topics, providing insights for future LLM deployment and culturally aware evaluation practices in multilingual contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。