arXiv:2506.11361cs.CLcs.CY2025-06

用道德情境测试大模型对不同人群的助人意愿偏见

The Biased Samaritan: LLM biases in Perceived Kindness

  • 让模型评估虚构人物助人意愿,量化性别种族年龄偏见
  • 发现基准群体为白人中青年男性,非基准群体更愿帮助
  • 可帮助开发者识别并修正模型输出中的隐性偏见

尽管大型语言模型(LLMs)已广泛应用于多个领域,但理解与缓解其偏见仍是持续挑战。本文提出一种新方法,用于评估多种生成式AI模型的群体偏见。通过提示模型判断虚构道德主体愿意主动干预的程度,我们定量分析了不同性别、种族和年龄群体在模型中的表现。本研究不同于以往工作之处在于,旨在确定商业模型的基准身份特征及其与其他群体的关系。我们试图判断这些偏见是正向、中性还是负向,以及其强度。分析结果显示:各模型均将白人中青年或年轻成年男性视为基准群体;然而,跨模型普遍趋势显示,非基准群体被模型认为更具助人意愿。该方法有效区分了常被混淆的两种偏见类型,有助于客观评估大模型中的偏见,并使用户或开发者能在输出或未来训练中加以考量。

原文摘要 · Abstract (English)

While Large Language Models (LLMs) have become ubiquitous in many fields, understanding and mitigating LLM biases is an ongoing issue. This paper provides a novel method for evaluating the demographic biases of various generative AI models. By prompting models to assess a moral patient's willingness to intervene constructively, we aim to quantitatively evaluate different LLMs' biases towards various genders, races, and ages. Our work differs from existing work by aiming to determine the baseline demographic identities for various commercial models and the relationship between the baseline and other demographics. We strive to understand if these biases are positive, neutral, or negative, and the strength of these biases. This paper can contribute to the objective assessment of bias in Large Language Models and give the user or developer the power to account for these biases in LLM output or in training future LLMs. Our analysis suggested two key findings: that models view the baseline demographic as a white middle-aged or young adult male; however, a general trend across models suggested that non-baseline demographics are more willing to help than the baseline. These methodologies allowed us to distinguish these two biases that are often tangled together.

大模型偏见道德评估社会偏见

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。