用姓名代替国籍标签,发现小模型偏见更严重且纠错能力差。
Obscured but Not Erased: Evaluating Nationality Bias in LLMs via Name-Based Bias Benchmarks
- 用文化相关姓名替代国籍标签,模拟真实应用场景测试模型偏见。
- 小模型准确率低117.7%,偏见得分高达9%(大模型仅3.5%)。
- 小模型在模糊情境下保留76%错误率,适合关注公平性的开发者参考。
大型语言模型(LLMs)即使在无显式人口统计标记时也可能存在对特定国籍的隐含偏见。本文提出一种基于姓名的基准评估方法,源自偏差问答基准(BBQ)数据集,研究将显式国籍标签替换为具有文化指示性的姓名后对模型偏见与准确性的影响。实验涵盖OpenAI、Google、Anthropic等公司的多款主流模型。结果显示,小模型准确率更低,偏见更显著:例如,在模糊语境下,Claude Haiku的刻板印象偏见得分为9%,而其大版本Claude Sonnet仅为3.5%,后者准确率高出117.7%。此外,小模型在模糊情境中仍保留较高错误率——如GPT-4o保留68%错误,GPT-4o-mini达76%,其他厂商模型亦呈现相似趋势。研究揭示了模型偏见的顽固性,强调其在全球化应用中的深远影响。
原文摘要 · Abstract (English)
Large Language Models (LLMs) can exhibit latent biases towards specific nationalities even when explicit demographic markers are not present. In this work, we introduce a novel name-based benchmarking approach derived from the Bias Benchmark for QA (BBQ) dataset to investigate the impact of substituting explicit nationality labels with culturally indicative names, a scenario more reflective of real-world LLM applications. Our novel approach examines how this substitution affects both bias magnitude and accuracy across a spectrum of LLMs from industry leaders such as OpenAI, Google, and Anthropic. Our experiments show that small models are less accurate and exhibit more bias compared to their larger counterparts. For instance, on our name-based dataset and in the ambiguous context (where the correct choice is not revealed), Claude Haiku exhibited the worst stereotypical bias scores of 9%, compared to only 3.5% for its larger counterpart, Claude Sonnet, where the latter also outperformed it by 117.7% in accuracy. Additionally, we find that small models retain a larger portion of existing errors in these ambiguous contexts. For example, after substituting names for explicit nationality references, GPT-4o retains 68% of the error rate versus 76% for GPT-4o-mini, with similar findings for other model providers, in the ambiguous context. Our research highlights the stubborn resilience of biases in LLMs, underscoring their profound implications for the development and deployment of AI systems in diverse, global contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。