arXiv:2502.10577cs.CLcs.AI2025-02中稿 · EMNLP被引 5

检测大模型对男性泛指语言的偏见,发现近三成回应仍显性别倾向。

Man Made Language Models? Evaluating LLMs' Perpetuation of Masculine Generics Bias

  • 用法语数据测试六款模型,对比有无男性泛指词时的回应差异。
  • 约27.57%的通用指令回复出现男性泛指偏见,含男性泛指词时高达78.55%。
  • 模型极少自发使用性别公平语言,提示需主动干预设计。

基于指令的大语言模型在上下文受限的提示下会传播甚至放大性别偏见。然而,对于通过性别化语言传递的上下文非受限(通用)指令中的偏见,尤其是男性泛指(MG)问题关注不足。男性泛指在诸多性别标记语言中被当作中性表达,用于指代混合性别群体或未知/非二元性别个体。但心理语言学研究证实,其并非中性,而是系统性引发性别偏见。本研究评估了本地与专有模型在法语环境下对通用指令的男性泛指偏见,分析其偏见率及性别公平语言(GFL)使用情况。我们从现有词源资源构建了一个16,000+条目的人类名词数据库,并在四个指令-响应数据集上,对六款模型在两种条件(含与不含男性泛指词)下进行评估。结果显示,约27.57%的模型在男性泛指过滤后的通用指令下仍出现偏见(含男性泛指词时达78.55%)。此外,模型极少自发采用性别公平语言。研究揭示大模型输出中男性泛指偏见持续存在,且对性别公平语言策略的倾向极低。

原文摘要 · Abstract (English)

Instruct-based large language models (LLMs) have been shown to propagate and even amplify gender bias when prompted with contextually constrained instructions (e.g., writing a text from a description or selecting a gendered pronoun). However, little attention has been paid to biases in responses to contextually unconstrained (generic) instructions conveyed by gendered language, particularly masculine generics (MG). MG, found in many gender-marked languages, denote the use of the masculine gender as a supposedly neutral reference to mixed-gender groups or individuals whose gender is unknown or non-binary. Yet, psycholinguistic studies demonstrate that MG are not neutral and systematically induce gender bias. This study investigates how both local and proprietary LLMs are MG-biased when responding to generic prompts in French, examining LLMs' MG bias rates and use of gender-fair language (GFL). We create a 16k+ human noun database from existing lexical resources and evaluate six LLMs on four instruction-response datasets under two conditions: prompts with and without MG. Overall, we find that $\approx$27.57% of LLMs' responses to MG-filtered generic instructions are MG-biased ($\approx$78.55% with MG-containing prompts). Moreover, we find that LLMs rarely use GFL spontaneously. These findings highlight the persistence of MG bias in LLM outputs and models' limited tendency towards GFL strategies.

性别偏见大模型语言公平法语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。