arXiv:2411.12074cs.CLcs.LG2024-11被引 2

提出新目标函数,有效降低上下文词嵌入中的性别偏见。

Mitigating Gender Bias in Contextual Word Embeddings

  • 设计针对掩码语言建模的新目标函数,缓解性别偏见。
  • 实验证明该方法在下游任务中保持性能不下降。
  • 揭示静态嵌入偏见主因是刻板印象姓名而非性别词汇。

词嵌入在众多自然语言处理任务中表现优异,但同时也捕获了社会中存在的刻板偏见,影响其在下游任务中的预测性能。尽管已有多种方法被提出并用于静态嵌入的去偏,但针对上下文嵌入的去偏研究仍十分有限。本文提出一种新的掩码语言建模(MLM)目标函数,显著缓解上下文嵌入中的性别偏见,同时保持下游任务性能。针对现有评估方法缺乏规范依据的问题,我们设计了更直接、符合去偏动机的新评估指标。此外,我们还提出改进静态嵌入的去偏方法,并通过大量实验与分析证明:静态嵌入偏见的主要来源是刻板印象姓名,而非性别词汇本身。所有实验与嵌入均基于英文,除非另有说明。

原文摘要 · Abstract (English)

Word embeddings have been shown to produce remarkable results in tackling a vast majority of NLP related tasks. Unfortunately, word embeddings also capture the stereotypical biases that are prevalent in society, affecting the predictive performance of the embeddings when used in downstream tasks. While various techniques have been proposed \cite{bolukbasi2016man, zhao2018learning} and criticized\cite{gonen2019lipstick} for static embeddings, very little work has focused on mitigating bias in contextual embeddings. In this paper, we propose a novel objective function for MLM(Masked-Language Modeling) which largely mitigates the gender bias in contextual embeddings and also preserves the performance for downstream tasks. Since previous works on measuring bias in contextual embeddings lack in normative reasoning, we also propose novel evaluation metrics that are straight-forward and aligned with our motivations in debiasing. We also propose new methods for debiasing static embeddings and provide empirical proof via extensive analysis and experiments, as to why the main source of bias in static embeddings stems from the presence of stereotypical names rather than gendered words themselves. All experiments and embeddings studied are in English, unless otherwise specified.\citep{bender2011achieving}.

词嵌入性别偏见去偏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。