arXiv:2602.21485cs.CLcs.HC2026-02

测试大模型对非裔美语的使用,发现存在明显偏差和刻板印象。

Evaluating the Usage of African-American Vernacular English in Large Language Models

  • 对比人类与大模型在非裔美语中的语法使用差异。
  • 模型普遍误用或少用典型非裔美语特征如'ain't'。
  • 模型生成内容复制了对非裔美国人的负面刻板印象。

在人工智能领域,自然语言理解任务的评估大多基于标准英语(如标准美式英语)。本文研究大型语言模型(LLMs)对非裔美语(AAVE)的表征准确性。通过分析来自《区域非裔美语语料库》(Corpus of Regional African American Language)和TwitterAAE的数据,识别出母语使用者在真实语境中使用AAVE语法特征(如'ain't')的典型模式。随后,我们向三款大模型发起生成任务,要求其以非裔美语生成文本,并与人类使用模式进行比较。结果显示,模型在多数情况下存在显著偏差:通常低估或误用非裔美语的语法特征。结合情感分析与人工检查,发现模型还重复了关于非裔美国人的刻板印象。这些发现凸显了训练数据多样性不足及需引入公平性方法以减少偏见传播的必要性。

原文摘要 · Abstract (English)

In AI, most evaluations of natural language understanding tasks are conducted in standardized dialects such as Standard American English (SAE). In this work, we investigate how accurately large language models (LLMs) represent African American Vernacular English (AAVE). We analyze three LLMs to compare their usage of AAVE to the usage of humans who natively speak AAVE. We first analyzed interviews from the Corpus of Regional African American Language and TwitterAAE to identify the typical contexts where people use AAVE grammatical features such as ain't. We then prompted the LLMs to produce text in AAVE and compared the model-generated text to human usage patterns. We find that, in many cases, there are substantial differences between AAVE usage in LLMs and humans: LLMs usually underuse and misuse grammatical features characteristic of AAVE. Furthermore, through sentiment analysis and manual inspection, we found that the models replicated stereotypes about African Americans. These results highlight the need for more diversity in training data and the incorporation of fairness methods to mitigate the perpetuation of stereotypes.

语言模型非裔美语偏见检测公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。