新指标发现翻译模型更倾向性别刻板印象,对女性指代常视而不见。
Are We Paying Attention to Her? Investigating Gender Disambiguation and Attention in Machine Translation
- 设计最小对比对准确率,检测模型是否真用性别线索
- 多数模型无视明确性别提示,偏信刻板印象性别分配
- 男性线索引发广泛注意,女性线索被忽略或弱响应
尽管现代神经机器翻译(NMT)中的性别偏见已受关注,但传统评估指标未能充分反映系统对上下文性别线索的利用程度。本文提出一种新评估指标——最小对比对准确率(MPA),衡量模型在仅通过性别代词差异区分的句子对中,是否依赖性别线索进行性别消歧。在英-意(EN--IT)语言对上评估多个NMT模型,结果显示,大多数情况下模型忽视可用的性别线索,转而采用统计上的刻板性别判断。在反刻板情境中,模型更一致地采纳男性线索,却常忽略女性线索。进一步分析编码器注意力头权重发现,虽然所有模型都一定程度编码性别信息,但男性线索引发更分散的响应,而女性线索则激发更集中、专一的反应。
原文摘要 · Abstract (English)
While gender bias in modern Neural Machine Translation (NMT) systems has received much attention, traditional evaluation metrics do not to fully capture the extent to which these systems integrate contextual gender cues. We propose a novel evaluation metric called Minimal Pair Accuracy (MPA), which measures the reliance of models on gender cues for gender disambiguation. MPA is designed to go beyond surface-level gender accuracy metrics by focusing on whether models adapt to gender cues in minimal pairs -- sentence pairs that differ solely in the gendered pronoun, namely the explicit indicator of the target's entity gender in the source language (EN). We evaluate a number of NMT models on the English-Italian (EN--IT) language pair using this metric, we show that they ignore available gender cues in most cases in favor of (statistical) stereotypical gender interpretation. We further show that in anti-stereotypical cases, these models tend to more consistently take masculine gender cues into account while ignoring the feminine cues. Furthermore, we analyze the attention head weights in the encoder component and show that while all models encode gender information to some extent, masculine cues elicit a more diffused response compared to the more concentrated and specialized responses to feminine gender cues.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。