分析翻译模型性别选择的触发因素,揭示其与人类认知的关联。
What Triggers my Model? Contrastive Explanations Inform Gender Choices by Translation Models
- 通过对比翻译生成显著性图,定位影响性别选择的源句词汇。
- 发现模型对性别词的敏感度与人类感知高度重合。
- 为缓解翻译模型性别偏见提供可解释性依据,适合关注AI公平性的研究者。
可解释性可用于理解神经机器翻译(NMT)或大语言模型(LLMs)等黑箱模型的决策过程。然而,现有研究在应对这些模型中的性别偏见问题上仍显不足。本文旨在超越单纯的偏见测量,探索其根源。基于性别模糊的自然源语言数据,本研究探究源句(英语)中哪些输入词元会触发目标语言(德语/西班牙语)翻译中的特定性别表达。为此,我们采用基于对比翻译的显著性归因方法,解决了缺乏评分阈值的挑战,并分析了源词在不同归因层级上对模型性别决策的影响。结果表明,模型归因与人类对性别的感知存在明显重叠。此外,我们还进行了显著词的语义分析。本研究展示了从性别角度理解模型翻译决策的重要性,揭示其与人类判断的一致性,并强调应利用此类信息以减轻性别偏见。
原文摘要 · Abstract (English)
Interpretability can be implemented to understand decisions taken by (black box) models, such as neural machine translation (NMT) or large language models (LLMs). Yet, research in this area has been limited in relation to a manifested problem in these models: gender bias. In this work, we aim to move away from simply measuring bias to exploring its origins. Working with gender-ambiguous natural source data, this exploratory study examines which context, in the form of input tokens in the source sentence (EN), influences (or triggers) the NMT model's choice of a certain gender inflection in the target languages (DE/ES). To analyse this, we compute saliency attribution based on contrastive translations. We first address the challenge of the lack of a scoring threshold and specifically examine different attribution levels of source words on the model's gender decisions in the translation. We compare salient source words with human perceptions of gender and demonstrate a noticeable overlap between human perceptions and model attribution. Additionally, we provide a linguistic analysis of salient words. Our work showcases the relevance of understanding model translation decisions in terms of gender, how this compares to human decisions and that this information should be leveraged to mitigate gender bias.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。