检测生成文本的模型对弱势群体存在系统性偏见,尤其误判非白人英语学习者。
Identifying Bias in Machine-generated Text Detection
- 基于学生作文数据集评估16个检测模型在四类属性上的偏差
- 非白人英语学习者作文被错误标记为机器生成的概率更高
- 经济劣势学生作文反而更少被误判,人类标注无显著偏见
随着文本生成能力的迅猛发展,机器生成文本检测技术也日益受到关注:即判断一段文本是模型生成还是人类撰写。尽管检测模型表现良好,但其可能带来重大负面影响。本文探究英语语境下机器生成文本检测系统的潜在偏见。我们构建了一个学生作文数据集,评估16种不同检测系统在性别、种族/族裔、英语学习者(ELL)身份和经济状况四个属性上的偏差。通过回归模型分析效应显著性与强度,并开展子群分析。结果表明,虽然各模型偏差表现不一致,但仍存在若干关键问题:多个模型倾向于将弱势群体判定为机器生成;英语学习者作文更易被误判为机器生成;经济劣势学生的作文较少被误判;而非白人英语学习者作文被错误标记为机器生成的比例显著高于白人同侪。最后,我们进行了人工标注,发现人类在检测任务中整体表现不佳,但在所研究属性上未表现出显著偏见。
原文摘要 · Abstract (English)
The meteoric rise in text generation capability has been accompanied by parallel growth in interest in machine-generated text detection: the capability to identify whether a given text was generated using a model or written by a person. While detection models show strong performance, they have the capacity to cause significant negative impacts. We explore potential biases in English machine-generated text detection systems. We curate a dataset of student essays and assess 16 different detection systems for bias across four attributes: gender, race/ethnicity, English-language learner (ELL) status, and economic status. We evaluate these attributes using regression-based models to determine the significance and power of the effects, as well as performing subgroup analysis. We find that while biases are generally inconsistent across systems, there are several key issues: several models tend to classify disadvantaged groups as machine-generated, ELL essays are more likely to be classified as machine-generated, economically disadvantaged students' essays are less likely to be classified as machine-generated, and non-White ELL essays are disproportionately classified as machine-generated relative to their White counterparts. Finally, we perform human annotation and find that while humans perform generally poorly at the detection task, they show no significant biases on the studied attributes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。