跨29种欧洲语言评估大模型性别刻板印象,发现女性多被关联‘美丽’‘共情’,男性则关联‘领导’‘坚强’。
EuroGEST: Investigating gender stereotypes in multilingual language models
- 基于翻译与语法规则扩展16类性别刻板印象至29种语言,构建多语言评测数据集。
- 24个模型均显示女性=‘美丽/共情/整洁’,男性=‘领导/强壮/专业’的强刻板倾向。
- 模型越大刻板印象越强,指令微调无法稳定降低偏见,适合公平性研究者使用。
大语言模型日益支持多语言,但多数性别偏见评测仍以英语为主。我们提出EuroGEST,一个涵盖英语和29种欧洲语言的多语言性别刻板印象评测数据集。该数据集在原有16类专家定义的性别刻板印象基础上,通过翻译工具、质量估计指标和形态学启发法进行扩展。人工评估证实,生成的翻译和性别标签在各语言中均具有高准确性。我们用EuroGEST评估了来自六个模型家族的24个多语言大模型,结果显示所有模型在所有语言中最强的刻板印象均为:女性是‘美丽’‘共情’‘整洁’,男性是‘领导’‘强壮’‘专业’。此外,更大的模型编码性别偏见更强烈,且指令微调并未一致地减少偏见。本工作强调了多语言公平性研究的必要性,并提供了可扩展的方法与资源,用于跨语言审计性别偏见。
原文摘要 · Abstract (English)
Large language models increasingly support multiple languages, yet most benchmarks for gender bias remain English-centric. We introduce EuroGEST, a dataset designed to measure gender-stereotypical reasoning in LLMs across English and 29 European languages. EuroGEST builds on an existing expert-informed benchmark covering 16 gender stereotypes, expanded in this work using translation tools, quality estimation metrics, and morphological heuristics. Human evaluations confirm that our data generation method results in high accuracy of both translations and gender labels across languages. We use EuroGEST to evaluate 24 multilingual language models from six model families, demonstrating that the strongest stereotypes in all models across all languages are that women are 'beautiful', 'empathetic' and 'neat' and men are 'leaders', 'strong, tough' and 'professional'. We also show that larger models encode gendered stereotypes more strongly and that instruction finetuning does not consistently reduce gendered stereotypes. Our work highlights the need for more multilingual studies of fairness in LLMs and offers scalable methods and resources to audit gender bias across languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。