针对无性别语言翻译中的性别偏见问题,提出新数据集与优化方法。
Translate With Care: Addressing Gender Bias, Neutrality, and Reasoning in Large Language Model Translations
- 构建TWC数据集,涵盖6种低中资源语言的3950个挑战场景。
- 多模型测试发现男性代词使用率高出女性4-6倍,尤其在职业语境中。
- 微调mBART-50可显著降低偏见,性能超越闭源大模型。
机器翻译在处理自然性别语言(如英语)与无性别语言(如波斯语、印尼语、芬兰语)之间的转换时,仍面临性别偏见与逻辑连贯性难题。本文提出Translate-with-Care(TWC)数据集,包含跨六种低至中资源语言的3,950个复杂翻译场景,用于评估翻译系统表现。对GPT-4、mBART-50、NLLB-200及Google Translate等技术的分析显示,所有模型在处理无性别语言时普遍存在性别刻板印象和推理错误,倾向于在可能受性别刻板印象影响的语境中选择男性代词。Google Translate与GPT-4尤为明显,在领导力与职业成功语境中,男性代词使用频率为女性的4至6倍。对mBART-50在TWC上进行微调后,显著缓解了上述偏见与错误,实现强泛化能力,且性能超越部分闭源大模型,同时保持开源。该研究强调需针对无性别语言设计专门的性别与语义一致性策略,推动更公平准确的翻译系统发展。
原文摘要 · Abstract (English)
Addressing gender bias and maintaining logical coherence in machine translation remains challenging, particularly when translating between natural gender languages, like English, and genderless languages, such as Persian, Indonesian, and Finnish. We introduce the Translate-with-Care (TWC) dataset, comprising 3,950 challenging scenarios across six low- to mid-resource languages, to assess translation systems' performance. Our analysis of diverse technologies, including GPT-4, mBART-50, NLLB-200, and Google Translate, reveals a universal struggle in translating genderless content, resulting in gender stereotyping and reasoning errors. All models preferred masculine pronouns when gender stereotypes could influence choices. Google Translate and GPT-4 showed particularly strong bias, favoring male pronouns 4-6 times more than feminine ones in leadership and professional success contexts. Fine-tuning mBART-50 on TWC substantially resolved these biases and errors, led to strong generalization, and surpassed proprietary LLMs while remaining open-source. This work emphasizes the need for targeted approaches to gender and semantic coherence in machine translation, particularly for genderless languages, contributing to more equitable and accurate translation systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。