arXiv:2410.10995cs.CL2024-10ACL被引 11

发现机器翻译质量评估存在性别偏见,男性化译文得分更高。

Watching the Watchers: Exposing Gender Disparities in Machine Translation Quality Estimation

  • 通过多语言多领域实验,发现未指明性别时男性化译文评分更高。
  • 女性指代词的正确识别率低于男性,上下文感知模型也更易出错。
  • 适合关注公平性、伦理和自动评估系统优化的研究者阅读。

质量评估(QE)在翻译流程中日益关键,涵盖数据筛选、训练与解码阶段。尽管现有QE指标已对齐人类判断,但其是否包含社会偏见尚未被充分关注。偏见的QE可能加剧特定群体的可见性与可用性差距。本文首次定义并检验了QE指标中的性别偏见及其对机器翻译的下游影响。在多个领域、数据集与语言上的实验表明:当源文本中人物性别未明确时,男性化译文得分显著高于女性化译文,而中性译文则被惩罚。即使上下文可推断性别,基于上下文的QE模型在识别女性指代时错误率仍高于男性。此外,带有偏见的QE会影响数据过滤与质量感知解码。研究呼吁重新聚焦于性别公平的QE指标开发与评估。

原文摘要 · Abstract (English)

Quality estimation (QE)-the automatic assessment of translation quality-has recently become crucial across several stages of the translation pipeline, from data curation to training and decoding. While QE metrics have been optimized to align with human judgments, whether they encode social biases has been largely overlooked. Biased QE risks favoring certain demographic groups over others, e.g., by exacerbating gaps in visibility and usability. This paper defines and investigates gender bias of QE metrics and discusses its downstream implications for machine translation (MT). Experiments with state-of-the-art QE metrics across multiple domains, datasets, and languages reveal significant bias. When a human entity's gender in the source is undisclosed, masculine-inflected translations score higher than feminine-inflected ones, and gender-neutral translations are penalized. Even when contextual cues disambiguate gender, using context-aware QE metrics leads to more errors in selecting the correct translation inflection for feminine referents than for masculine ones. Moreover, a biased QE metric affects data filtering and quality-aware decoding. Our findings underscore the need for a renewed focus on developing and evaluating QE metrics centered on gender.

质量评估性别偏见机器翻译

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。