arXiv:2607.28528cs.CL2026-07

AI语言模型正在强化主流英语霸权,边缘化非主流英语变体。

AI systems and the reproduction of (standard) language ideologies in World Englishes

  • 分析AI训练数据、设计与评价体系如何复制主流英语规范
  • 揭示全球北方对全球南方英语使用者的语法审查现象
  • 呼吁包容多元英语的算法设计,避免语言不公

大型语言模型的快速发展重新引发社会语言学和世界英语研究中的核心问题:谁定义合法英语?谁的英语被质疑?本文通过实证研究、媒体评论、社交媒体讨论及AI输出案例,揭示人工智能系统在训练数据、设计规范、评估基准、用户反馈与公众话语中,持续复制并强化以内部圈层英语为标准的语言意识形态,边缘化非主导英语变体。文章以AI生成语料中对'delve'一词的过度关注为例,说明全球北方如何对全球南方英语使用者实施语言规范监控。同时指出,尽管AI可能通过偏好标准形式而同质化英语,但其海量语料与全球南方用户的标注工作也带来英语多样性的新机遇。论文主张,生成式AI正成为语言意识形态再生产的新场域,必须推动更具包容性的设计,承认英语的多样性,以应对因语言合法性不平等带来的现实危害。

原文摘要 · Abstract (English)

The rapid growth of large language models (LLMs) has resurrected age-old questions in sociolinguistics and world Englishes, such as who decides what counts as legitimate English, whose English is suspect etc. This paper examines how AI systems, their uses and discourse on them reflect, reinforce, and occasionally challenge (standard) language ideologies, which privilege Inner Circle norms and marginalize non-dominant Englishes. Drawing on evidence from empirical studies, media commentary, social media debates, and examples from AI outputs, the paper shows that AI technologies reproduce dominant language ideologies at different levels: training data, design protocols, evaluation benchmarks, user feedback and public commentary. The analysis uses the public controversy over AI-sounding language, especially the fixation on the word delve, to illustrate how speakers of English from the Global North police the English language norms of Global South English users. The paper also identifies what Christian Mair has called a "standardisation paradox": AI may homogenize English by privileging standard forms and at the same time pluralize Englishes through exposure to wide-ranging corpora and annotation work carried out by Global South users. In doing so, the paper argues that generative AI is reigniting long-standing debates in World Englishes about standardization, legitimacy, and the ownership of English, now playing out in algorithmic systems, model training, evaluation practices, and public discourse, where non-dominant Englishes are increasingly conflated with AI-generated speech. Discussing AI systems as a site where language ideologies are (re)produced, the paper argues for more inclusive design approaches that recognize the plurality of Englishes in order to address the real-world negative consequences of treating some as more legitimate than others.

语言意识形态AI伦理世界英语公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。