arXiv:2510.06841cs.CL2025-10中稿 · publication at the…被引 3

构建大规模数据集,检测机器翻译质量评估中的性别偏见。

GAMBIT+: A Challenge Set for Evaluating Gender Bias in Machine Translation Quality Estimation Metrics

  • 基于性别模糊职业词,构建跨语言平行数据集
  • 33个语种对中,同文本仅性别词不同,评估得分差异
  • 适合研究公平性、质量评估与多语言偏见的学者使用

机器翻译中的性别偏见已广为人知,但自动质量评估(QE)指标中的偏见仍鲜有研究。现有分析受限于小规模数据、职业覆盖窄及语言种类少。为此,我们提出GAMBIT+挑战集,专门用于评估QE指标在处理性别模糊职业术语时的表现。基于英文性别模糊职业语料库GAMBIT,扩展至三种源语言(无性别/自然性别),并涵盖十一种具有语法性别特征的目标语言,形成33个源-目标语言对。每条源文本对应两个仅在职业词语法性别上不同的目标版本(阳性 vs. 阴性),所有相关语法成分均相应调整。理想情况下,无偏的QE指标应对两版本评分相近。该数据集规模大、覆盖广、结构完全平行,使职业层级与跨语言系统性对比成为可能,为细粒度偏见分析提供支持。

原文摘要 · Abstract (English)

Gender bias in machine translation (MT) systems has been extensively documented, but bias in automatic quality estimation (QE) metrics remains comparatively underexplored. Existing studies suggest that QE metrics can also exhibit gender bias, yet most analyses are limited by small datasets, narrow occupational coverage, and restricted language variety. To address this gap, we introduce a large-scale challenge set specifically designed to probe the behavior of QE metrics when evaluating translations containing gender-ambiguous occupational terms. Building on the GAMBIT corpus of English texts with gender-ambiguous occupations, we extend coverage to three source languages that are genderless or natural-gendered, and eleven target languages with grammatical gender, resulting in 33 source-target language pairs. Each source text is paired with two target versions differing only in the grammatical gender of the occupational term(s) (masculine vs. feminine), with all dependent grammatical elements adjusted accordingly. An unbiased QE metric should assign equal or near-equal scores to both versions. The dataset's scale, breadth, and fully parallel design, where the same set of texts is aligned across all languages, enables fine-grained bias analysis by occupation and systematic comparisons across languages.

性别偏见质量评估多语言数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。