arXiv:2601.16349cs.CLcs.AI2026-01被引 1

测试10个大模型对不同地区的偏见,发现差距极大。

Regional Bias in Large Language Models

  • 用100个中立情景提问,测模型对地区的选择偏好。
  • GPT-3.5偏见最严重(得分9.5),Claude 3.5最弱(2.5)。
  • 为检测和缓解地理偏见提供可量化的评估框架。

本研究探讨大型语言模型(LLMs)中的区域偏见问题,这是人工智能公平性与全球代表性的重要议题。我们使用100个精心设计的提示,评估了十种主流大模型:GPT-3.5、GPT-4o、Gemini 1.5 Flash、Gemini 1.0 Pro、Claude 3 Opus、Claude 3.5 Sonnet、Llama 3、Gemma 7B、Mistral 7B 和 Vicuna-13B。这些提示在语境中立的情景下引导模型做出地区选择。我们提出一种基于提示的评估框架 FAZE,以10分制衡量区域偏见,分数越高表示越倾向于特定地区。实验结果表明各模型间偏见程度差异显著,其中 GPT-3.5 的偏见得分最高(9.5),Claude 3.5 Sonnet 最低(2.5)。这说明区域偏见会严重影响大模型在真实跨文化应用中的可靠性、公平性和包容性。本工作推动了人工智能公平性研究,强调了构建包容性评估框架与系统化方法的重要性,以识别并减轻语言模型中的地理偏见。

原文摘要 · Abstract (English)

This study investigates regional bias in large language models (LLMs), an emerging concern in AI fairness and global representation. We evaluate ten prominent LLMs: GPT-3.5, GPT-4o, Gemini 1.5 Flash, Gemini 1.0 Pro, Claude 3 Opus, Claude 3.5 Sonnet, Llama 3, Gemma 7B, Mistral 7B, and Vicuna-13B using a dataset of 100 carefully designed prompts that probe forced-choice decisions between regions under contextually neutral scenarios. We introduce FAZE, a prompt-based evaluation framework that measures regional bias on a 10-point scale, where higher scores indicate a stronger tendency to favor specific regions. Experimental results reveal substantial variation in bias levels across models, with GPT-3.5 exhibiting the highest bias score (9.5) and Claude 3.5 Sonnet scoring the lowest (2.5). These findings indicate that regional bias can meaningfully undermine the reliability, fairness, and inclusivity of LLM outputs in real-world, cross-cultural applications. This work contributes to AI fairness research by highlighting the importance of inclusive evaluation frameworks and systematic approaches for identifying and mitigating geographic biases in language models.

大模型偏见检测公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。