现有性别偏见评估基准各有局限,综合使用更准确
Blind Men and the Elephant: Diverse Perspectives on Gender Stereotypes in Benchmark Datasets
- 用社会心理学方法平衡数据,提升评估一致性
- 简单平衡后,不同测评方法相关性显著提高
- 适合研究语言模型偏见检测与改进的学者参考
准确衡量语言模型中的性别刻板印象偏见是一项复杂任务,现有基准测试未能全面反映这一多维度挑战。本文分析了内在刻板印象基准间的不一致现象,指出当前基准仅捕捉了性别刻板印象的部分特征,孤立使用时呈现的是片面视角。以StereoSet和CrowS-Pairs为例,研究发现数据分布显著影响测评结果。通过引入社会心理学框架对两类基准的数据在性别刻板印象各维度进行平衡,实验表明,即使采用简单的平衡策略,也能显著提升不同测量方法之间的相关性。研究强调了语言模型中性别刻板印象的复杂性,并为开发更精细的偏见检测与缓解技术指明新方向。
原文摘要 · Abstract (English)
Accurately measuring gender stereotypical bias in language models is a complex task with many hidden aspects. Current benchmarks have underestimated this multifaceted challenge and failed to capture the full extent of the problem. This paper examines the inconsistencies between intrinsic stereotype benchmarks. We propose that currently available benchmarks each capture only partial facets of gender stereotypes, and when considered in isolation, they provide just a fragmented view of the broader landscape of bias in language models. Using StereoSet and CrowS-Pairs as case studies, we investigated how data distribution affects benchmark results. By applying a framework from social psychology to balance the data of these benchmarks across various components of gender stereotypes, we demonstrated that even simple balancing techniques can significantly improve the correlation between different measurement approaches. Our findings underscore the complexity of gender stereotyping in language models and point to new directions for developing more refined techniques to detect and reduce bias.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。