arXiv:2608.30884cs.CLcs.AI2026-08中稿 · EMNLP

构建德英双语数据集,评估语言模型对跨文化酷儿偏见的复制情况。

Evaluating and Mitigating Anti-LGBTQ Biases in German and Multilingual Language Models

  • 融合德语酷儿群体自述刻板印象与德文版WinoQueer构建评测数据集。
  • 8个模型均存在反酷儿偏见,不同身份和模型间表现差异显著。
  • 社区内容微调可平均降低偏见,但效果不一,适合多元文化研究者参考。

尽管性别与种族偏见在语言模型中已广受关注,针对非英语语境下的反酷儿偏见研究仍严重不足。现有基准常忽略文化与语言差异,且依赖性别表征。本文提出一个德英双语多语言基准数据集,结合德语区酷儿群体提供的刻板印象与WinoQueer的德文翻译。该数据用于评估八种不同规模与架构的语言模型,并探索通过社区媒体及进步媒体内容微调以缓解偏见。结果显示,语言模型普遍再现反酷儿刻板印象,且在不同身份与模型间存在显著差异。翻译数据与社区数据的差异凸显了多语言偏见评估中文化适配的重要性。微调虽平均降低偏见,但在不同模型与身份上效果不一致。警告:本文本包含反酷儿仇恨言论与刻板印象。

原文摘要 · Abstract (English)

While gender and racial biases in language models have been widely studied, anti-LGBTQ biases remain underexplored, particularly beyond English. Existing benchmarks often do not capture cultural and linguistic variation and rely on gender representations. This paper introduces a multilingual German-English benchmark dataset for the evaluation of anti-LGBTQ biases in language models. It combines community-sourced stereotypes from German-speaking queer individuals with a German translation of WinoQueer. The data is used to evaluate eight language models across sizes and architectures and explore mitigation through fine-tuning on community and progressive media content. Results show that language models reproduce anti-queer stereotypes, with variation across identities and models. Differences between the translated and community-based data highlight the importance of cultural adaptation for multilingual bias evaluation. Fine-tuning reduces bias on average, but not consistently across models and identities. Warning: This text contains examples of anti-queer hateful language and stereotypes.

偏见评估多语言酷儿研究模型微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。