arXiv:2507.18338cs.CL2025-07Conference of the …被引 2

用不确定性评估机器翻译中的性别偏见,发现高准确率不等于合理不确定。

Uncertainty Quantification for Evaluating Machine Translation Bias

  • 用语义不确定性衡量翻译中性别推断的合理性
  • 高准确率翻译仍可能缺乏适当不确定性
  • 适用于有歧义和无歧义场景,适合研究偏见的学者

机器翻译模型的预测不确定性通常用作质量估计的代理。本文认为,当输入存在歧义时,模型不仅应能自信翻译,还应保持不确定性。我们利用不确定性来衡量翻译系统中的性别偏见:当源句包含未明确标记性别的词,而目标语言需指定性别时,模型需从上下文推断性别,易受偏见影响。以往工作通过性别准确率衡量偏见,但无法处理歧义情况。本文使用语义不确定性,能够评估歧义与非歧义句子的偏见表现,发现高翻译准确率与适当不确定性无相关性,且去偏策略对两类情况影响不同。

原文摘要 · Abstract (English)

The predictive uncertainty of machine translation (MT) models is typically used as a quality estimation proxy. In this work, we posit that apart from confidently translating when a single correct translation exists, models should also maintain uncertainty when the input is ambiguous. We use uncertainty to measure gender bias in MT systems. When the source sentence includes a lexeme whose gender is not overtly marked, but whose target-language equivalent requires gender specification, the model must infer the appropriate gender from the context and can be susceptible to biases. Prior work measured bias via gender accuracy, however it cannot be applied to ambiguous cases. Using semantic uncertainty, we are able to assess bias when translating both ambiguous and unambiguous source sentences, and find that high translation accuracy does not correlate with exhibiting uncertainty appropriately, and that debiasing affects the two cases differently.

机器翻译偏见评估不确定性量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。