arXiv:2605.07622cs.CL2026-05

分析荷兰语BERT如何学习性别信息,发现模型难以响应明确的女性提示。

Is She Even Relevant? When BERT Ignores Explicit Gender Cues

论文配图:Is She Even Relevant? When BERT Ignores Explicit Gender Cues
图 1 · 摘自论文原文
  • 用线性SVM追踪训练中性别信息的编码演变过程
  • 第20轮后性别可被线性区分,但女性提示仍被忽略
  • 通用形式默认男性化,反刻板印象仍预测不准

大语言模型中的性别偏见主要集中在英语,而具有语法或形态性别标记的语言研究较少。本文对从零训练的荷兰语BERT模型进行检查点级分析,探究性别信息在结合明显形态性别标记与通用形式的语言中的出现时机与方式。通过提取训练全程的上下文嵌入,利用线性SVM构建动态性别子空间,追踪性别编码的形成与演化。我们测试在控制句式中(如 'Zij is een loodgieter')明确性别线索能否覆盖习得的统计关联(如 'plumber' → male)。结果挑战了上下文嵌入能稳健整合线索的假设:尽管性别在第20轮后已清晰线性可分且分布在多个维度,模型仍无法有效更新内部性别表示。刻板印象的职业-性别配对预测远优于反刻板印象,荷兰语通用形式系统性默认为男性,即使上下文明确指向女性。表明该模型在探测的性别方向上,上下文适应性不足,显性性别线索未被可靠反映,导致持续的男性默认行为。

原文摘要 · Abstract (English)

Gender bias in large language models has primarily been investigated for English, while languages with grammatical or morphological gender remain comparatively understudied. This paper investigates how and when gender information emerges in a Dutch BERT model trained from scratch, offering one of the first checkpoint-level analyses of bias formation in a Transformer architecture for a language combining overt morphological gender marking and generic forms. By extracting contextual embeddings throughout training, we construct dynamic gender subspaces using linear SVMs to trace when gender becomes linearly encoded and how this encoding evolves over time. Contextual embeddings are often assumed to integrate contextual cues robustly, allowing models to adjust the representation of a word depending on its more local usage. We therefore test whether explicit gender cues in controlled sentence templates (e.g., Zij is een loodgieter ('She is a plumber')) can override learned statistical associations (plumber -> male). Our findings challenge this assumption: although gender becomes clearly linearly separable around epoch 20 and is distributed across multiple embedding dimensions, the model struggles to update its internal gender representation in light of explicit contextual cues in short sentence templates. Stereotypical gender-profession pairings are predicted far more accurately than anti-stereotypical ones, and generic forms in Dutch systematically default to a male interpretation, even when the context explicitly denotes a female referent. Together, our results seem to indicate that contextualization in the representations learned by our Dutch BERT model is not sufficiently dynamic along the probed gender direction: explicit gender cues in anti-stereotypical contexts are not reliably reflected in the resulting representations, resulting in persistent male-default behaviour.

性别偏见BERT语言模型荷兰语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。