构建首个系统性评估AI文本检测器偏见的基准测试框架
BAID: A Benchmark for Bias Assessment of AI Detectors
- 设计覆盖7类社会语言因素的20万+样本数据集
- 发现主流检测器对少数群体文本召回率显著偏低
- 适合教育科技、AI伦理领域研究者参考
AI生成文本检测器在教育和职业场景中日益普及。已有研究揭示了其在英语学习者群体中的孤立偏见,但缺乏对更广泛社会语言因素的系统评估。本文提出BAID,一个全面评估AI检测器偏见的框架。作为该框架的一部分,我们构建了超过20万条样本,涵盖七大类别:人口统计、年龄、学段、方言、正式程度、政治倾向和主题。我们还通过精心设计的提示生成每条样本的合成版本,保持原内容的同时体现特定子群体的写作风格。利用该数据集,我们评估了四种开源前沿AI文本检测器,发现检测性能存在持续差异,尤其对代表性不足群体的文本召回率偏低。本研究贡献了一个可扩展、透明的AI检测器审计方法,强调在公开部署前进行偏见敏感评估的重要性。
原文摘要 · Abstract (English)
AI-generated text detectors have recently gained adoption in educational and professional contexts. Prior research has uncovered isolated cases of bias, particularly against English Language Learners (ELLs) however, there is a lack of systematic evaluation of such systems across broader sociolinguistic factors. In this work, we propose BAID, a comprehensive evaluation framework for AI detectors across various types of biases. As a part of the framework, we introduce over 200k samples spanning 7 major categories: demographics, age, educational grade level, dialect, formality, political leaning, and topic. We also generated synthetic versions of each sample with carefully crafted prompts to preserve the original content while reflecting subgroup-specific writing styles. Using this, we evaluate four open-source state-of-the-art AI text detectors and find consistent disparities in detection performance, particularly low recall rates for texts from underrepresented groups. Our contributions provide a scalable, transparent approach for auditing AI detectors and emphasize the need for bias-aware evaluation before these tools are deployed for public use.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。