提出双去偏算法,在消除性别刻板印象的同时保留真实性别信息。
Dual Debiasing: Remove Stereotypes and Keep Factual Gender for Fair Language Modeling and Translation
- 通过双重去偏机制,分离并消除刻板印象中的性别关联。
- 在英文语言模型中显著降低性别偏差,且保持真实性别信息不变。
- 适合需要公平性与准确性兼顾的文本生成与翻译任务。
减轻语言模型对性别刻板印象的依赖,是构建可靠、可用语言技术的关键挑战。去偏的核心在于确保模型保留其多样化的语言能力,包括完成语言任务及公平表征各类性别的能力。为此,本文提出基于模型适配的双去偏算法(2DAMA)。该方法能有效减少英语语言模型中的性别偏见,并首次实现翻译过程中刻板倾向的缓解。其关键优势在于保留了语言模型中编码的真实性别信息,这些信息在多种自然语言处理任务中具有重要价值。
原文摘要 · Abstract (English)
Mitigation of biases, such as language models' reliance on gender stereotypes, is a crucial endeavor required for the creation of reliable and useful language technology. The crucial aspect of debiasing is to ensure that the models preserve their versatile capabilities, including their ability to solve language tasks and equitably represent various genders. To address this issue, we introduce a streamlined Dual Dabiasing Algorithm through Model Adaptation (2DAMA). Novel Dual Debiasing enables robust reduction of stereotypical bias while preserving desired factual gender information encoded by language models. We show that 2DAMA effectively reduces gender bias in English and is one of the first approaches facilitating the mitigation of stereotypical tendencies in translation. The proposed method's key advantage is the preservation of factual gender cues, which are useful in a wide range of natural language processing tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。