发现大模型对不同性别和代词的道德判断存在系统性偏差。
Widespread Gender and Pronoun Bias in Moral Judgments Across LLMs
- 通过改写句子中的代词和性别标记,生成1.485万条语义等价句测试模型
- 第三人称单数句更常被判定为公平,第二人称则遭惩罚;非二元性别受青睐,男性受贬低
- 揭示训练数据与对齐机制导致的隐性偏见,提醒警惕道德评估中的模型风险
大型语言模型(LLMs)日益用于评估道德或伦理陈述,但其判断可能反映社会与语言偏见。本研究基于ETHICS数据集中的550个平衡基础句,每句生成26个反事实变体,系统改变代词与人口统计标记,共得到14,850条语义等价句子。评估了六类模型家族(Grok、GPT、LLaMA、Gemma、DeepSeek、Mistral),使用统计公平差异(SPD)测量公平性判断与组间差异。结果显示显著偏见:第三人称单数句更常被判定为“公平”,第二人称则被惩罚;性别标记影响最强,非二元性别主体始终更受青睐,男性主体则被贬低。我们推测这些模式源于训练过程中的分布与对齐偏见,强调在道德应用中需采取针对性公平干预。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used to assess moral or ethical statements, yet their judgments may reflect social and linguistic biases. This work presents a controlled, sentence-level study of how grammatical person, number, and gender markers influence LLM moral classifications of fairness. Starting from 550 balanced base sentences from the ETHICS dataset, we generated 26 counterfactual variants per item, systematically varying pronouns and demographic markers to yield 14,850 semantically equivalent sentences. We evaluated six model families (Grok, GPT, LLaMA, Gemma, DeepSeek, and Mistral), and measured fairness judgments and inter-group disparities using Statistical Parity Difference (SPD). Results show statistically significant biases: sentences written in the singular form and third person are more often judged as "fair'', while those in the second person are penalized. Gender markers produce the strongest effects, with non-binary subjects consistently favored and male subjects disfavored. We conjecture that these patterns reflect distributional and alignment biases learned during training, emphasizing the need for targeted fairness interventions in moral LLM applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。