揭示预训练语料中非裔美国人语言的严重缺失与偏见
Data Caricatures: On the Representation of African American Language in Pretraining Corpora
- 分析12个英文语料库,量化非裔美国人语言占比
- AAL仅占0.007%至0.18%,远低于人口比例
- 多数自动过滤器更倾向保留主流英语,加剧语言不公
通过定量实验、人工评估和定性分析,我们评估了12个主要英文开源预训练语料库中非裔美国人语言(AAL)的代表性。重点关注其来源、变体和自然性,反映非裔美国人使用者群体的真实表达。研究发现,所有语料库中AAL占比均低于美国人口比例,最低仅占0.007%,最高为0.18%。超过25%的C4语料中AAL文本可能被判定为不适合大模型生成,且易强化有害刻板印象。此外,多数自动化过滤机制更倾向于保留白人主流英语(WME)文本,而非AAL,进一步加剧语言代表性失衡。
原文摘要 · Abstract (English)
With a combination of quantitative experiments, human judgments, and qualitative analyses, we evaluate the quantity and quality of African American Language (AAL) representation in 12 predominantly English, open-source pretraining corpora. We specifically focus on the sources, variation, and naturalness of included AAL texts representing the AAL-speaking community. We find that AAL is underrepresented in all evaluated pretraining corpora compared to US demographics, constituting as few as 0.007% and at most 0.18% of documents. We also find that more than 25% of AAL texts in C4 may be perceived as inappropriate for LLMs to generate and to reinforce harmful stereotypes. Finally, we find that most automated filters are more likely to conserve White Mainstream English (WME) texts over AAL in pretraining corpora.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。