arXiv:2507.16442cs.CL2025-07中稿 · RANLP 2025 data an…被引 4

构建荷兰语偏见检测数据集,揭示语言模型中的文化偏见差异

Dutch CrowS-Pairs: Adapting a Challenge Dataset for Measuring Social Biases in Language Models for Dutch

  • 基于CrowS-Pairs构建1463对荷兰语句子,覆盖9类社会偏见
  • 发现荷兰语模型偏见程度低于英语模型,但角色设定会改变偏见水平
  • 首次系统评估荷兰语模型偏见,适用于多语言公平性研究

语言模型易放大不公平和有害的刻板印象。尽管近年来对模型偏见测量的关注增加,但多数研究集中于英语。本文引入首个针对荷兰语的、基于美国原版CrowS-Pairs的数据集,包含1463个句子对,覆盖性取向、性别、残疾等9类偏见。这些句子对由涉及弱势群体与优势群体的对比句构成。实验表明,BERTje、RobBERT、多语言BERT、GEITje和Mistral-7B等模型在各类偏见中均表现出显著偏差。对比英法模型(使用英文和法文版本的CrowS-Pairs)发现,英语模型偏见最严重,而荷兰语模型偏见最少。此外,给模型赋予特定人格角色会显著改变其表现的偏见水平。结果表明,偏见程度受语言与语境影响,文化与语言因素在塑造模型偏见中起关键作用。

原文摘要 · Abstract (English)

Warning: This paper contains explicit statements of offensive stereotypes which might be upsetting. Language models are prone to exhibiting biases, further amplifying unfair and harmful stereotypes. Given the fast-growing popularity and wide application of these models, it is necessary to ensure safe and fair language models. As of recent considerable attention has been paid to measuring bias in language models, yet the majority of studies have focused only on English language. A Dutch version of the US-specific CrowS-Pairs dataset for measuring bias in Dutch language models is introduced. The resulting dataset consists of 1463 sentence pairs that cover bias in 9 categories, such as Sexual orientation, Gender and Disability. The sentence pairs are composed of contrasting sentences, where one of the sentences concerns disadvantaged groups and the other advantaged groups. Using the Dutch CrowS-Pairs dataset, we show that various language models, BERTje, RobBERT, multilingual BERT, GEITje and Mistral-7B exhibit substantial bias across the various bias categories. Using the English and French versions of the CrowS-Pairs dataset, bias was evaluated in English (BERT and RoBERTa) and French (FlauBERT and CamemBERT) language models, and it was shown that English models exhibit the most bias, whereas Dutch models the least amount of bias. Additionally, results also indicate that assigning a persona to a language model changes the level of bias it exhibits. These findings highlight the variability of bias across languages and contexts, suggesting that cultural and linguistic factors play a significant role in shaping model biases.

偏见检测语言模型荷兰语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。