发现大模型默认偏好美式英语,揭示语言偏见的深层根源
Which English Do LLMs Prefer? Triangulating Structural Bias Towards American English in Foundation Models
- 构建1813组英美变体语料,提出无需训练的动态对齐方法DiAlign
- 三阶段验证:预训练数据、分词成本、生成输出均显示美式英语占优
- 首次系统揭示大模型在英式英语上的结构性偏见,适合关注AI公平性的研究者
大型语言模型(LLMs)在高风险领域日益普及,但仅支持有限的语言设置,尤其是美式英语(AmE),尽管英语具有全球多样性和殖民历史。本文从后殖民视角出发,探讨数据收集的地缘政治历史、数字主导地位和语言标准化如何塑造大模型开发流程。聚焦两种主流标准变体——美式英语(AmE)与英式英语(BrE),我们构建了包含1,813个变体对的精选语料库,并提出DiAlign方法,一种基于分布证据的无训练动态对齐技术。通过三个阶段的三角验证:(i)对六大数据预训练语料库的审计显示系统性倾向美式英语;(ii)分词分析表明英式形式分段成本更高;(iii)生成评估显示模型输出持续偏好美式英语。这是首个系统且多维度考察大模型开发全流程中标准英语变体不对称性的研究。结果表明当代大模型将美式英语作为事实上的标准,引发语言同质化、知识不公与全球人工智能部署不平等的担忧,同时为构建更包容方言的大语言技术提供实践方向。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly deployed in high-stakes domains, yet they expose only limited language settings, most notably "English (US)," despite the global diversity and colonial history of English. Through a postcolonial framing to explain the broader significance, we investigate how geopolitical histories of data curation, digital dominance, and linguistic standardization shape the LLM development pipeline. Focusing on two dominant standard varieties, American English (AmE) and British English (BrE), we construct a curated corpus of 1,813 AmE--BrE variants and introduce DiAlign, a dynamic, training-free method for estimating dialectal alignment using distributional evidence. We operationalize structural bias by triangulating evidence across three stages: (i) audits of six major pretraining corpora reveal systematic skew toward AmE, (ii) tokenizer analyses show that BrE forms incur higher segmentation costs, and (iii) generative evaluations show a persistent AmE preference in model outputs. To our knowledge, this is the first systematic and multi-faceted examination of dialectal asymmetries in standard English varieties across the phases of LLM development. We find that contemporary LLMs privilege AmE as the de facto norm, raising concerns about linguistic homogenization, epistemic injustice, and inequity in global AI deployment, while motivating practical steps toward more dialectally inclusive language technologies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。