arXiv:2501.04662cs.CL2025-01NAACL被引 26

揭示语言模型文化偏见根源,发现阿拉伯语中高频实体易出错

On The Origin of Cultural Biases in Language Models: From Pre-training Data to Linguistic Phenomena

  • 构建阿拉伯-英语双语数据集CAMeL-2,分析文化实体在预训练中的表现
  • 模型在阿拉伯语中对高频多义实体识别准确率显著下降,尤其当词义重叠其他使用阿拉伯字母的语言时
  • 指出频率分词机制加剧偏见,大词汇量模型问题更严重,适合关注公平性研究者阅读

语言模型在非西方语言中常表现出对西方文化实体的强烈偏好。本文通过分析预训练数据中实体表征及跨语言语言现象差异,探究此类文化偏见的成因。我们提出CAMeL-2,一个包含58,086个阿拉伯与西方文化相关实体、367个带掩码自然上下文的阿拉伯-英语平行基准。评估显示,模型在英文测试中文化间性能差距较小,但在阿拉伯语中明显增大。模型在阿拉伯语中对预训练高频出现且具有多重词义的实体表现不佳,且当实体与非阿拉伯语但使用阿拉伯文字母的语言存在高词汇重叠时问题更严重。我们进一步证明频率分词机制导致此问题,且随着阿拉伯语词汇量增大而恶化。数据集将开源。

原文摘要 · Abstract (English)

Language Models (LMs) have been shown to exhibit a strong preference towards entities associated with Western culture when operating in non-Western languages. In this paper, we aim to uncover the origins of entity-related cultural biases in LMs by analyzing several contributing factors, including the representation of entities in pre-training data and the impact of variations in linguistic phenomena across languages. We introduce CAMeL-2, a parallel Arabic-English benchmark of 58,086 entities associated with Arab and Western cultures and 367 masked natural contexts for entities. Our evaluations using CAMeL-2 reveal reduced performance gaps between cultures by LMs when tested in English compared to Arabic. We find that LMs struggle in Arabic with entities that appear at high frequencies in pre-training, where entities can hold multiple word senses. This also extends to entities that exhibit high lexical overlap with languages that are not Arabic but use the Arabic script. Further, we show how frequency-based tokenization leads to this issue in LMs, which gets worse with larger Arabic vocabularies. We will make CAMeL-2 available at: https://github.com/tareknaous/camel2

文化偏见语言模型阿拉伯语公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。