arXiv:2506.10491cs.CL2025-06被引 12

研究发现语言模型在不同任务中表现出明显偏见,尤其在薪资谈判建议中。

Surface Fairness, Deep Bias: A Comparative Study of Bias in Language Models

  • 通过改写任务形式,发现模型对用户身份有显著偏见。
  • 在薪资谈判场景中,模型给出的建议存在明显性别与社会背景偏见。
  • 适合关注模型伦理、公平性及个性化应用的研究者阅读。

现代语言模型基于海量数据训练,这些数据不可避免地包含涉及性别、族裔、年龄等的争议性与刻板内容,导致模型在表达观点或生成结果时产生偏见。本文探讨了大语言模型(LLMs)中偏见的各种代理度量方式。实验发现,在多主题基准测试(MMLU)上使用预设人格提示时,模型得分差异微小且基本随机。然而,若将任务改为让模型评价用户的回答,则显示出更显著的偏见。更重要的是,当要求模型提供薪资谈判建议时,偏见表现尤为突出。随着大语言模型向具备记忆与个性化能力发展,用户无需主动设定角色,模型已能自动推断其社会人口学特征,这使得偏见问题从新角度浮现。

原文摘要 · Abstract (English)

Modern language models are trained on large amounts of data. These data inevitably include controversial and stereotypical content, which contains all sorts of biases related to gender, origin, age, etc. As a result, the models express biased points of view or produce different results based on the assigned personality or the personality of the user. In this paper, we investigate various proxy measures of bias in large language models (LLMs). We find that evaluating models with pre-prompted personae on a multi-subject benchmark (MMLU) leads to negligible and mostly random differences in scores. However, if we reformulate the task and ask a model to grade the user's answer, this shows more significant signs of bias. Finally, if we ask the model for salary negotiation advice, we see pronounced bias in the answers. With the recent trend for LLM assistant memory and personalization, these problems open up from a different angle: modern LLM users do not need to pre-prompt the description of their persona since the model already knows their socio-demographics.

语言模型偏见检测伦理风险

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。