arXiv:2606.30152cs.CLcs.AI2026-06

首次解耦上下文嵌入中的语法性别与语义偏见,提升模型公平性。

Estimating Grammatical Gender Directions in Contextual Embeddings under Controlled and Natural Contexts

论文配图:Estimating Grammatical Gender Directions in Contextual Embeddings under Controlled and Natural Contexts
图 1 · 摘自论文原文
  • 构建受控与自然语境数据集,分离语法性别信号。
  • 中心点估计器在去除性别泄露上表现最优,优于传统方法。
  • 适合关注语言模型偏见、性别平等的研究者使用。

上下文语言模型在西班牙语等性别标记语言中会混淆语法性别与社会语义偏见。现有去偏方法仅针对静态词向量,未探索上下文表示中的双维性别解耦。为此,我们首次尝试对上下文嵌入进行语法性别与语义污染的分离。构建了受控模板和自然维基百科语境下的平衡数据集,包含无生命名词,并设计融合中心点、支持向量机(SVM)与线性判别分析(LDA)的性别方向估计算法,以及感知污染的加权策略。提出一套双目标评估指标,以平衡无生命名词的语法性别泄漏抑制与职业词的语义性别区分保留。结果表明,未经加权的受控语境产生最纯净的语法性别方向,中心点估计算法性能优于判别基线。

原文摘要 · Abstract (English)

Contextual language models conflate grammatical gender and social semantic bias in gendered languages such as Spanish. Existing gender debiasing approaches only operate on static word embeddings leaving contextual representations unexplored for this two dimensional gender disentanglement. To address the this issue, we make the first attempt to disentangle grammatical gender from semantic contamination for contextual embeddings. We construct both controlled templates and natural Wikipedia contexts to build balanced datasets of inanimate nouns, and design a framework equipped with centroid, Support Vector Machine (SVM) and Linear Discriminant Analysis (LDA) gender direction estimators as well as contamination-aware weighting strategies. A set of dual-objective evaluation metrics is proposed to balance the suppression of grammatical gender leakage on inanimate nouns and the preservation of semantic gender distinctions for occupation terms. The results reveal that unweighted controlled contexts yield the purest grammatical gender direction, and the centroid estimator achieves better performance than discriminative baselines.

性别偏见上下文嵌入去偏语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。