arXiv:2503.04372cs.CL2025-03EMNLP被引 5

量化机器翻译对性别模糊职业词的系统性偏见

Assumed Identities: Quantifying Gender Bias in Machine Translation of Gender-Ambiguous Occupational Terms

  • 用概率聚合方法评估翻译中隐含的性别倾向
  • 发现多国语言翻译普遍偏离真实职业性别分布
  • 适合研究算法公平性与社会偏见的学者

机器翻译系统在处理性别模糊的职业术语时,需在无明确语境下进行性别指派。尽管单个翻译未必带偏见,但系统性地将特定职业与特定性别关联,可能反映并强化社会刻板印象。传统基于单一标准答案的评估方式无法应对此类问题。为此,我们提出GRAPE——一种基于概率的评估指标,用于分析模型输出的聚合模式;同时构建GAMBIT数据集,包含英文中性别模糊的职业术语。利用GRAPE,我们评估了多个MT系统,并考察其在希腊语和法语中的性别化翻译是否符合社会刻板印象、现实职业性别分布及规范标准。

原文摘要 · Abstract (English)

Machine Translation (MT) systems frequently encounter gender-ambiguous occupational terms, where they must assign gender without explicit contextual cues. While individual translations in such cases may not be inherently biased, systematic patterns-such as consistently translating certain professions with specific genders-can emerge, reflecting and perpetuating societal stereotypes. This ambiguity challenges traditional instance-level single-answer evaluation approaches, as no single gold standard translation exists. To address this, we introduce GRAPE, a probability-based metric designed to evaluate gender bias by analyzing aggregated model responses. Alongside this, we present GAMBIT, a benchmarking dataset in English with gender-ambiguous occupational terms. Using GRAPE, we evaluate several MT systems and examine whether their gendered translations in Greek and French align with or diverge from societal stereotypes, real-world occupational gender distributions, and normative standards

机器翻译性别偏见评估方法公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。