为无性别语言巴斯克语构建性别偏见评测基准
Gender Bias in MT for a Genderless Language: New Benchmarks for Basque
- 用巴斯克语职业词翻译成有性别语言,测试性别偏见
- 多数模型倾向选男性形式,部分模型男性译文质量略高
- 适合研究语言公平性与低资源语言的学者参考
大型语言模型和机器翻译系统在日常中广泛应用,但其输出可能复制训练数据中的性别偏见。现有评估资源多针对英语,难以适用于其他语言。本文针对低资源且无性别特征的巴斯克语,提出两个新数据集:WinoMTeus 将 WinoMT 基准适配至巴斯克语,考察性别中立的职业在翻译为西班牙语、法语等有性别语言时的表现;FLORES+Gender 扩展 FLORES 基准,评估从西班牙语、英语等有性别语言翻译至巴斯克语时,译文质量是否随指代对象性别而异。我们评估了多个通用 LLM 及开源与专有翻译系统,结果显示系统普遍倾向使用男性形式,部分模型对男性指代者的翻译质量略高。总体表明,性别偏见仍深植于这些模型中,强调需发展兼顾语言特征与文化背景的评估方法。
原文摘要 · Abstract (English)
Large language models (LLMs) and machine translation (MT) systems are increasingly used in our daily lives, but their outputs can reproduce gender bias present in the training data. Most resources for evaluating such biases are designed for English and reflect its sociocultural context, which limits their applicability to other languages. This work addresses this gap by introducing two new datasets to evaluate gender bias in translations involving Basque, a low-resource and genderless language. WinoMTeus adapts the WinoMT benchmark to examine how gender-neutral Basque occupations are translated into gendered languages such as Spanish and French. FLORES+Gender, in turn, extends the FLORES+ benchmark to assess whether translation quality varies when translating from gendered languages (Spanish and English) into Basque depending on the gender of the referent. We evaluate several general-purpose LLMs and open and proprietary MT systems. The results reveal a systematic preference for masculine forms and, in some models, a slightly higher quality for masculine referents. Overall, these findings show that gender bias is still deeply rooted in these models, and highlight the need to develop evaluation methods that consider both linguistic features and cultural context.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。