arXiv:2410.09992cs.CL2024-10EMNLP被引 23

测试大模型在道德判断中的性别偏见,发现多数模型存在显著偏差。

Evaluating Gender Bias of LLMs in Making Morality Judgements

  • 构建平行故事数据集GenMO,对比男女角色的道德评价差异。
  • GPT-3.5-turbo在24%样本中出现偏见,部分模型对女性角色偏好达85%。
  • 揭示模型参数与偏见关系,适合关注AI伦理的研究者参考。

大型语言模型(LLMs)在自然语言处理任务中表现出色,但仍存在社会偏见,尤其是性别偏见。本文研究当前闭源与开源模型在道德判断中的性别偏见问题。为此,我们构建并引入新数据集GenMO,包含成对短故事,分别以男性和女性角色为主角。测试了GPT家族(GPT-3.5-turbo、GPT-3.5-turbo-instruct、GPT-4-turbo)、Llama 3及3.1系列(8B/70B)、Mistral-7B以及Claude 3系列(Sonnet和Opus)。令人惊讶的是,尽管经过安全审查,所有生产级模型均表现出显著性别偏见,其中GPT-3.5-turbo在24%样本中给出有偏见的判断。此外,所有模型一致倾向于女性角色,GPT在68%-85%情况下呈现偏见,Llama 3在约81%-85%实例中如此。研究还分析了模型参数对偏见的影响,并探讨了现实场景中模型在道德决策中暴露偏见的情况。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have shown remarkable capabilities in a multitude of Natural Language Processing (NLP) tasks. However, these models are still not immune to limitations such as social biases, especially gender bias. This work investigates whether current closed and open-source LLMs possess gender bias, especially when asked to give moral opinions. To evaluate these models, we curate and introduce a new dataset GenMO (Gender-bias in Morality Opinions) comprising parallel short stories featuring male and female characters respectively. Specifically, we test models from the GPT family (GPT-3.5-turbo, GPT-3.5-turbo-instruct, GPT-4-turbo), Llama 3 and 3.1 families (8B/70B), Mistral-7B and Claude 3 families (Sonnet and Opus). Surprisingly, despite employing safety checks, all production-standard models we tested display significant gender bias with GPT-3.5-turbo giving biased opinions in 24% of the samples. Additionally, all models consistently favour female characters, with GPT showing bias in 68-85% of cases and Llama 3 in around 81-85% instances. Additionally, our study investigates the impact of model parameters on gender bias and explores real-world situations where LLMs reveal biases in moral decision-making.

大模型性别偏见道德判断AI伦理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。