arXiv:2412.07303cs.CL2024-12中稿 · presentation at Th…被引 9

构建菲律宾语偏见评估基准,揭示低资源语言模型中的性别与性向偏见。

Filipino Benchmarks for Measuring Sexist and Homophobic Bias in Multilingual Language Models from Southeast Asia

  • 基于英语偏见数据集文化适配,生成7074组菲律宾语挑战对。
  • 多语言模型在菲律宾语中存在显著性别与反同性恋偏见。
  • 模型偏见程度与其训练语料量正相关,适合关注多元文化偏见的研究者。

现有针对多语言模型的偏见研究已证实,在高资源语言上存在性别刻板印象。本文扩展该研究,提出菲律宾语版CrowS-Pairs和WinoQueer基准,用于评估预训练语言模型(PLMs)在菲律宾语——一种低资源东南亚语言——中的性别歧视与反同性恋偏见。这些基准由7,074个经文化适配的新挑战对构成,其构建过程被详细记录,以指导未来类似工作。我们在掩码与因果多语言模型(包括基于东南亚数据预训练的模型)上应用这些基准,发现模型存在显著偏见。同时,模型在特定语言上的偏见程度受其预训练语料中该语言数据量的影响。本研究提供的基准与洞见可为未来分析和缓解多语言模型偏见提供基础。

原文摘要 · Abstract (English)

Bias studies on multilingual models confirm the presence of gender-related stereotypes in masked models processing languages with high NLP resources. We expand on this line of research by introducing Filipino CrowS-Pairs and Filipino WinoQueer: benchmarks that assess both sexist and anti-queer biases in pretrained language models (PLMs) handling texts in Filipino, a low-resource language from the Philippines. The benchmarks consist of 7,074 new challenge pairs resulting from our cultural adaptation of English bias evaluation datasets, a process that we document in detail to guide similar forthcoming efforts. We apply the Filipino benchmarks on masked and causal multilingual models, including those pretrained on Southeast Asian data, and find that they contain considerable amounts of bias. We also find that for multilingual models, the extent of bias learned for a particular language is influenced by how much pretraining data in that language a model was exposed to. Our benchmarks and insights can serve as a foundation for future work analyzing and mitigating bias in multilingual models.

偏见评估多语言模型菲律宾语社会偏见

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。