arXiv:2410.12499cs.CL2024-10被引 4

测试开源大模型在性别、宗教和种族上的偏见,发现系统性不公。

With a Grain of SALT: Are LLMs Fair Across Social Dimensions?

  • 用SALT数据集测试多个模型在五类场景中的偏见表现。
  • 不同群体在辩论中胜率差异明显,职业建议中负面角色分配更集中。
  • 提出自动化评估方法并验证结果,适合关注AI公平性的研究者。

本文系统分析了开源大语言模型在性别、宗教和种族维度上的偏见。研究使用SALT(社会适切性在大模型生成文本中的体现)数据集,评估Llama和Gemma等小规模模型,该数据集包含五类偏见触发任务:通用辩论、立场辩论、职业建议、问题解决和简历生成。通过测量通用辩论中的胜率及立场辩论中负面角色的分配情况来量化偏见。针对职业建议、问题解决和简历生成等真实应用场景,对输出进行匿名化处理,并采用DeepSeek-R1作为自动化评估器。同时识别并缓解基于LLM的评估中存在的评价偏差、位置偏差和长度偏差,并通过人工评估验证结果。研究发现,不同模型间存在持续的极化现象,某些社会群体受到系统性优待或不利对待。通过引入SALT,本文构建了一个全面的偏见分析基准,强调了开发公平人工智能系统所需的强大偏见缓解策略。

原文摘要 · Abstract (English)

This paper presents a systematic analysis of biases in open-source Large Language Models (LLMs), across gender, religion, and race. Our study evaluates bias in smaller-scale Llama and Gemma models using the SALT ($\textbf{S}$ocial $\textbf{A}$ppropriateness in $\textbf{L}$LM-Generated $\textbf{T}$ext) dataset, which incorporates five distinct bias triggers: General Debate, Positioned Debate, Career Advice, Problem Solving, and CV Generation. To quantify bias, we measure win rates in General Debate and the assignment of negative roles in Positioned Debate. For real-world use cases, such as Career Advice, Problem Solving, and CV Generation, we anonymize the outputs to remove explicit demographic identifiers and use DeepSeek-R1 as an automated evaluator. We also address inherent biases in LLM-based evaluation, including evaluation bias, positional bias, and length bias, and validate our results through human evaluations. Our findings reveal consistent polarization across models, with certain demographic groups receiving systematically favorable or unfavorable treatment. By introducing SALT, we provide a comprehensive benchmark for bias analysis and underscore the need for robust bias mitigation strategies in the development of equitable AI systems.

大模型偏见检测公平性SALT

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。