评测主流大模型在职业与犯罪场景中的偏见,发现去偏举措可能引发新不公平。
LLM Bias Evaluation: Gender, Racial, and Age Disparities in Occupational and Crime Scenarios
- 对比4个2024年主流模型在职业与犯罪场景的性别、种族、年龄偏差
- 职业场景女性出现率比美国劳工数据高37%,犯罪场景性别偏差达54%
- 去偏措施反而可能加剧特定群体占比,揭示去偏悖论
大语言模型(LLMs)在高风险决策中应用日益广泛,其偏见评估至关重要。本文对2024年发布的四个领先模型——Gemini 1.5 Pro、Llama 3 70B、Claude 3 Opus和GPT-4o——进行了全面评估,分析其在职业与犯罪场景中的性别、种族和年龄偏见。研究发现,模型在职业场景中常过度描绘女性角色,相较于美国劳工统计局(US BLS)数据存在37%的偏差;在犯罪场景中,性别偏差为54%,种族偏差为28%,年龄偏差为17%。值得注意的是,旨在减少性别与种族偏见的努力往往导致某些子群体被过度代表,形成新的不公平现象,即‘去偏悖论’,揭示当前去偏技术的局限性,凸显亟需更有效的公平性提升方法。
原文摘要 · Abstract (English)
LLM bias evaluation is critical as large language models (LLMs) increasingly influence high-stakes decisions. This paper provides a comprehensive assessment of gender, racial, and age disparities in leading LLMs, revealing that debiasing efforts often create new fairness trade-offs. Recent advancements in LLMs have been notable, yet widespread enterprise adoption remains limited due to various constraints. This paper examines bias in LLMs - a crucial issue affecting their usability, reliability, and fairness. Our study evaluates gender bias in occupational scenarios and gender, age, and racial bias in crime scenarios across four leading LLMs released in 2024: Gemini 1.5 Pro, Llama 3 70B, Claude 3 Opus, and GPT-4o. Findings reveal that LLMs often depict female characters more frequently than male ones in various occupations, showing a 37% deviation from US BLS data. In crime scenarios, deviations from US FBI data are 54% for gender, 28% for race, and 17% for age. Critically, we observe that efforts to reduce gender and racial bias often lead to outcomes that may over-index one sub-class, potentially exacerbating disparities - a "debiasing paradox" that highlights the limitations of current bias mitigation techniques and underscores the need for more effective approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。