测试大模型性别偏见,发现其输出偏向真实统计数据中的固有偏见。
Assessing Gender Bias in LLMs: Comparing LLM Outputs with Human Perceptions and Official Statistics
- 构建新评测集,避免训练数据泄露
- 五款模型均偏离性别中立,倾向统计偏差
- 适合关注模型公平性与社会影响的研究者
本研究通过对比大语言模型(LLMs)的性别认知与人类受访者、美国劳工统计局数据以及50%无偏基准,评估其性别偏见。我们基于职业数据和角色特定句子构建了新的评估集,该集未包含于常见训练数据中,有效防止数据泄露与测试污染。五款LLMs被用于单字回答预测各角色性别。采用Kullback-Leibler(KL)散度比较模型输出与人类感知、统计数据及中立基准。所有模型均显著偏离性别中立,且更接近统计资料,反映出其内在偏见。
原文摘要 · Abstract (English)
This study investigates gender bias in large language models (LLMs) by comparing their gender perception to that of human respondents, U.S. Bureau of Labor Statistics data, and a 50% no-bias benchmark. We created a new evaluation set using occupational data and role-specific sentences. Unlike common benchmarks included in LLM training data, our set is newly developed, preventing data leakage and test set contamination. Five LLMs were tested to predict the gender for each role using single-word answers. We used Kullback-Leibler (KL) divergence to compare model outputs with human perceptions, statistical data, and the 50% neutrality benchmark. All LLMs showed significant deviation from gender neutrality and aligned more with statistical data, still reflecting inherent biases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。