arXiv:2511.03880cs.CLcs.CY2025-11

分析三种低资源语言翻译数据集的性别偏见,揭示数据量大不等于质量高。

Evaluating Machine Translation Datasets for Low-Web Data Languages: A Gendered Lens

  • 从性别视角评估阿法·奥罗莫语等三语言翻译数据集
  • 男性角色占比超七成,女性常被刻板描绘且存在有害内容
  • 适合关注NLP公平性与低资源语言研究者阅读

随着低资源语言在自然语言处理研究中的增多,大规模数据集的构建日益重要。然而,过度追求数据量而忽视质量,可能导致语言技术性能差,并加剧社会偏见。本文聚焦阿法·奥罗莫语、阿姆哈拉语和提格里尼亚语三种低资源语言的机器翻译数据集,重点考察其中的性别表征。研究发现,训练数据以政治和宗教文本为主,而基准数据集则集中于新闻、健康与体育领域;数据中男性角色占主导,姓名、动词语法性别及文本描述均呈现显著性别偏差,且存在对女性的有害与攻击性表述,尤其在数据量最大的语言中更为明显。这表明数据规模不能保证质量。本研究呼吁对低资源语言数据集进行更深入审视,并尽早干预有害内容。

原文摘要 · Abstract (English)

As low-resourced languages are increasingly incorporated into NLP research, there is an emphasis on collecting large-scale datasets. But in prioritizing quantity over quality, we risk 1) building language technologies that perform poorly for these languages and 2) producing harmful content that perpetuates societal biases. In this paper, we investigate the quality of Machine Translation (MT) datasets for three low-resourced languages--Afan Oromo, Amharic, and Tigrinya, with a focus on the gender representation in the datasets. Our findings demonstrate that while training data has a large representation of political and religious domain text, benchmark datasets are focused on news, health, and sports. We also found a large skew towards the male gender--in names of persons, the grammatical gender of verbs, and in stereotypical depictions in the datasets. Further, we found harmful and toxic depictions against women, which were more prominent for the language with the largest amount of data, underscoring that quantity does not guarantee quality. We hope that our work inspires further inquiry into the datasets collected for low-resourced languages and prompts early mitigation of harmful content. WARNING: This paper contains discussion of NSFW content that some may find disturbing.

机器翻译性别偏见低资源语言数据质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。