通过相似性网络分析英译意中的性别偏见,助力实现更公平的机器翻译。
Identifying Gender Stereotypes and Biases in Automated Translation from English to Italian using Similarity Networks
- 基于人权法与语言学定义性别偏见,聚焦 she/lei 与 he/lui 等关键词。
- 计算目标词间余弦相似度,揭示模型对性别关联的隐含认知偏差。
- 为训练无性别偏见的翻译模型提供可量化依据,适合政策与AI伦理研究者。
本文是语言学、法学与计算机科学的跨学科合作,旨在评估自动化翻译系统中的性别刻板印象与偏见。我们主张采用无性别翻译以促进性别包容性并提升机器翻译的客观性。研究聚焦于英译意场景中的性别偏见识别。首先,依据人权法与语言学文献定义性别偏见;随后,识别 she/lei 与 he/lui 等性别特定词汇作为关键元素。通过计算这些目标词与其他词汇间的余弦相似度,揭示模型对语义关系的感知。利用数值特征,有效评估偏见的强度与方向。研究结果为开发和训练无性别偏见的翻译算法提供了可量化的洞察。
原文摘要 · Abstract (English)
This paper is a collaborative effort between Linguistics, Law, and Computer Science to evaluate stereotypes and biases in automated translation systems. We advocate gender-neutral translation as a means to promote gender inclusion and improve the objectivity of machine translation. Our approach focuses on identifying gender bias in English-to-Italian translations. First, we define gender bias following human rights law and linguistics literature. Then we proceed by identifying gender-specific terms such as she/lei and he/lui as key elements. We then evaluate the cosine similarity between these target terms and others in the dataset to reveal the model's perception of semantic relations. Using numerical features, we effectively evaluate the intensity and direction of the bias. Our findings provide tangible insights for developing and training gender-neutral translation algorithms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。