评测大模型性别偏见,发现其在招聘等场景中存在刻板印象和不公。
GenderBench: Evaluation Suite for Gender Biases in LLMs
- 设计14个探针,量化19种性别有害行为
- 测试12个大模型,均现刻板印象与不平等倾向
- 开源工具包,适合研究者评估模型公平性
我们提出GenderBench——一个全面评估大模型性别偏见的评测套件。该套件包含14个探针,用于量化大模型表现出的19种与性别相关的有害行为。我们公开发布GenderBench为开源可扩展库,以提升领域内评测的可复现性和稳健性。同时公布了对12个大模型的评估结果。测量显示,这些模型在处理刻板印象推理、生成文本中的性别平等表现上存在系统性缺陷,且在高风险场景(如招聘)中偶发歧视性行为。
原文摘要 · Abstract (English)
We present GenderBench -- a comprehensive evaluation suite designed to measure gender biases in LLMs. GenderBench includes 14 probes that quantify 19 gender-related harmful behaviors exhibited by LLMs. We release GenderBench as an open-source and extensible library to improve the reproducibility and robustness of benchmarking across the field. We also publish our evaluation of 12 LLMs. Our measurements reveal consistent patterns in their behavior. We show that LLMs struggle with stereotypical reasoning, equitable gender representation in generated texts, and occasionally also with discriminatory behavior in high-stakes scenarios, such as hiring.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。