arXiv:2601.20731cs.CLcs.AI2026-01Conference of the …被引 1

LLM在句子补全中再现性别规范偏见,对酷儿群体输出更负面。

QueerGen: How LLMs Reflect Societal Norms on Gender and Sexuality in Sentence Completion Tasks

  • 通过句子补全测试,量化模型对不同性别标记群体的响应差异。
  • 掩码语言模型对酷儿标记个体产生最负面情感与更高毒性内容。
  • 闭源自回归模型反而对普通个体输出更有害内容,偏见形式因模型而异。

本文研究大型语言模型(LLMs)如何复制社会规范,特别是异性恋顺性规范,并考察这些规范在文本生成中的可测量偏见。我们分析了在三种主体类别——酷儿标记、非酷儿标记和标准化的‘无标记’类别——中,明确的性别或性取向信息是否影响模型回应。代表性不平衡通过四个维度进行量化:情感倾向、尊重程度、毒性水平和预测多样性。结果显示,掩码语言模型(MLMs)对酷儿标记主体产生最不利的情感、更高的毒性及更负面的评价;自回归语言模型(ARLMs)部分缓解此类模式,而闭源ARLMs则对无标记主体产生更具伤害性的输出。结果表明,LLMs复现了规范化的社会假设,但偏见的形式与程度强烈依赖于具体模型特征,可能重新分配而非消除代表性伤害。

原文摘要 · Abstract (English)

This paper examines how Large Language Models (LLMs) reproduce societal norms, particularly heterocisnormativity, and how these norms translate into measurable biases in their text generations. We investigate whether explicit information about a subject's gender or sexuality influences LLM responses across three subject categories: queer-marked, non-queer-marked, and the normalized "unmarked" category. Representational imbalances are operationalized as measurable differences in English sentence completions across four dimensions: sentiment, regard, toxicity, and prediction diversity. Our findings show that Masked Language Models (MLMs) produce the least favorable sentiment, higher toxicity, and more negative regard for queer-marked subjects. Autoregressive Language Models (ARLMs) partially mitigate these patterns, while closed-access ARLMs tend to produce more harmful outputs for unmarked subjects. Results suggest that LLMs reproduce normative social assumptions, though the form and degree of bias depend strongly on specific model characteristics, which may redistribute, but not eliminate, representational harms.

大模型偏见性别规范语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。