16个大模型普遍将性别等同于生理特征,忽视跨性别者身份。
Gender Trouble in Language Models: An Empirical Audit Guided by Gender Performativity Theory
- 用性别表演理论分析模型如何建构性别
- 大模型强化性别与生理的绑定关系
- 适合关注公平性与社会影响的研究者
语言模型会固化并传播有害的性别刻板印象。现有方法多聚焦于消除职业等中性词与'女人''男人'等性别词之间的关联,但这种做法仍流于表面,因为偏见不仅体现在关联上,更源于性别本身的建构。基于性别表演理论,我们发现语言模型常将性别视为与生物性别绑定的二元结构,导致跨性别及非二元性别身份被抹除或病理化。通过对16种不同架构、训练数据集和规模的语言模型进行实证测试,结果表明:模型普遍将性别等同于生理特征;非二元性别表达往往被系统性忽略或贬低。此外,尽管更大模型在性能基准上表现更优,却也表现出更强的性别-生理绑定,进一步固化狭窄的性别观。研究呼吁重新定义语言模型中的性别偏见,推动更深层的公平性治理。
原文摘要 · Abstract (English)
Language models encode and subsequently perpetuate harmful gendered stereotypes. Research has succeeded in mitigating some of these harms, e.g. by dissociating non-gendered terms such as occupations from gendered terms such as 'woman' and 'man'. This approach, however, remains superficial given that associations are only one form of prejudice through which gendered harms arise. Critical scholarship on gender, such as gender performativity theory, emphasizes how harms often arise from the construction of gender itself, such as conflating gender with biological sex. In language models, these issues could lead to the erasure of transgender and gender diverse identities and cause harms in downstream applications, from misgendering users to misdiagnosing patients based on wrong assumptions about their anatomy. For FAccT research on gendered harms to go beyond superficial linguistic associations, we advocate for a broader definition of 'gender bias' in language models. We operationalize insights on the construction of gender through language from gender studies literature and then empirically test how 16 language models of different architectures, training datasets, and model sizes encode gender. We find that language models tend to encode gender as a binary category tied to biological sex, and that gendered terms that do not neatly fall into one of these binary categories are erased and pathologized. Finally, we show that larger models, which achieve better results on performance benchmarks, learn stronger associations between gender and sex, further reinforcing a narrow understanding of gender. Our findings lead us to call for a re-evaluation of how gendered harms in language models are defined and addressed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。