arXiv:2609.02651cs.CL2026-09

构建荷兰语反酷儿偏见评估数据集,发现模型对跨性别者偏见严重。

WinoQueer-NL: Assessing Bias in Dutch Language Models toward LGBTQ+ Identities

论文配图:WinoQueer-NL: Assessing Bias in Dutch Language Models toward LGBTQ+ Identities
图 1 · 摘自论文原文
  • 基于英文WinoQueer改造荷兰语版本,融合文化语境设计句对
  • 42,906条语料中跨性别相关句型被模型支持达97%(最高偏差)
  • 适用于研究荷兰语模型偏见或伦理安全的开发者与政策制定者

尽管英语语言模型已被广泛研究反酷儿偏见,但荷兰语模型仍缺乏相关评估。为此,我们基于英文WinoQueer基准,构建了文化和语言适配的荷兰语数据集,包含成对的刻板印象与反刻板印象句子。通过面向43名荷兰酷儿参与者的在线调查,验证了171个刻板印象中的145个具有文化相关性,并从自由文本中识别出22个新偏见。最终发布的数据集含42,906条句子,使用多种荷兰语专用及多语言模型(包括掩码语言模型MLMs和自回归语言模型ARLMs)进行评估,以对比刻板与反刻板句子的对数似然得分衡量偏见。尽管整体平均偏见得分接近中性(~50%),但细粒度分析显示:部分模型对跨性别身份相关句型的支持率高达97%,而对同性恋相关句型仅6%;跨性别与非二元身份始终呈现最高偏见得分。结果强调了基于文化的评估数据集对检测和缓解荷兰语模型中针对边缘群体偏见的重要性。

原文摘要 · Abstract (English)

While English language models have been widely examined for anti-queer bias, Dutch models remain understudied. To address this gap, we developed a culturally and linguistically adapted Dutch dataset based on the English WinoQueer benchmark, containing pairs of stereotypical and counter-stereotypical sentences. To validate and expand it, we conducted an online survey with 43 Dutch queer participants, confirming 145 of 171 stereotypes as culturally relevant and identifying 22 new biases through free-text responses. The final released dataset, comprising 42,906 sentences, was evaluated using a range of Dutch-specific and multilingual models, including both masked language models (MLMs) and autoregressive language models (ARLMs), with bias measured via a score comparing log-likelihoods of stereotypical versus counter-stereotypical sentences. While the mean bias score across models appeared neutral (~50%), closer analysis revealed significant disparities: some models favored stereotypical sentences up to 97% of the time for transgender identities, but only 6% of the time for gay-related pairs, with transgender and non-binary identities consistently receiving the highest bias scores. Our findings highlight the importance of culturally grounded datasets for evaluating and mitigating biases that disproportionately impact marginalized groups in Dutch language models.

语言模型偏见评估荷兰语酷儿研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。