arXiv:2504.03352cs.CLcs.CY2025-04EMNLP被引 2

区分刻板印象与反刻板印象,构建更精准的检测基准

StereoDetect: Detecting Stereotypes and Anti-stereotypes the Correct Way Using Social Psychological Underpinnings

  • 基于社会心理学提出五元定义,厘清概念边界
  • 发现小模型和GPT-4o常误判反刻板印象
  • 提供对齐定义的高质量数据集,适合负责任AI研究

刻板印象危害巨大,其检测至关重要。但现有研究多聚焦于刻板偏见,未能清晰区分刻板印象、反刻板印象、刻板偏见与一般偏见,严重制约该领域进展。本研究基于社会心理学构建概念框架,提出五元定义与精确术语,以解决这一难题。我们识别出现有基准在该任务中的关键缺陷,并开发了名为StereoDetect的高质量、定义对齐的基准数据集。实验表明,子100亿参数语言模型及GPT-4o普遍存在反刻板印象误分类问题,且无法识别中性过度概括。通过定性与定量对比验证,StereoDetect在模型评估上优于现有基准。数据集与代码已开源。

原文摘要 · Abstract (English)

Stereotypes are known to have very harmful effects, making their detection critically important. However, current research predominantly focuses on detecting and evaluating stereotypical biases, thereby leaving the study of stereotypes in its early stages. Our study revealed that many works have failed to clearly distinguish between stereotypes and stereotypical biases, which has significantly slowed progress in advancing research in this area. Stereotype and Anti-stereotype detection is a problem that requires social knowledge; hence, it is one of the most difficult areas in Responsible AI. This work investigates this task, where we propose a five-tuple definition and provide precise terminologies disentangling stereotypes, anti-stereotypes, stereotypical bias, and general bias. We provide a conceptual framework grounded in social psychology for reliable detection. We identify key shortcomings in existing benchmarks for this task of stereotype and anti-stereotype detection. To address these gaps, we developed StereoDetect, a well curated, definition-aligned benchmark dataset designed for this task. We show that sub-10B language models and GPT-4o frequently misclassify anti-stereotypes and fail to recognize neutral overgeneralizations. We demonstrate StereoDetect's effectiveness through multiple qualitative and quantitative comparisons with existing benchmarks and models fine-tuned on them. The dataset and code is available at https://github.com/KaustubhShejole/StereoDetect.

刻板印象负责任AI文本检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。