语言模型在各个阶段都对方言存在系统性偏见,导致性能下降。
The Dialect Tax: Dialectal Biases Persist throughout the Language Modeling Pipeline
- 用平行方言语料追踪模型从分词到推理的全过程偏差。
- 方言文本在预训练中引发更大梯度差异,学习难度高于无关标准英语。
- 即使改用字符级分词,方言性能差距依然存在,说明问题贯穿全链路。
语言模型在方言上的系统性性能差距已广为人知,但其在现代语言建模流程中的根源尚不明确。本研究通过保持语义一致、仅改变表层形式的平行英语方言语料,追踪了这一“方言税”在整个自然语言处理流程中的演变。我们首先确认语言模型能将匹配的标准美式英语(SAE)与方言文本视为语义等价。然而,下游性能差距对应着深层表示差异:在不同模型家族和代际中,现代语言模型在分词、预训练、后训练和推理阶段均不对等编码方言文本。令人震惊的是,采用字符级反事实分词器虽消除了输入输出不对称,却无法消除方言准确率差距。预训练阶段,方言对的梯度更新差异大于完全无关的SAE文档对,表明模型更难从语义等价的方言内容中学习。后训练阶段,奖励模型表现出上下文依赖的不稳定方言偏好:孤立的非标准英语(AAVE)词汇得分更高,但在完整推理上下文中则出现任务与模型依赖的方言惩罚。总体而言,方言税并非由单一环节造成,而是贯穿语言建模全过程的累积效应。
原文摘要 · Abstract (English)
Systematic dialectal performance gaps in language models (LMs) are well documented, but the source of these disparities within the modern language modeling pipeline remains unclear. Our study traces this "dialect tax" across the natural language processing pipeline. Using parallel English dialect corpora that hold meaning fixed while varying surface form, we first confirm that LMs recognize matched Standard American English (SAE) and dialectal texts as semantically equivalent. However, we discover further representational gaps corresponding to downstream performance gaps. Across model families and generations, modern LMs still encode dialectal texts unequally during tokenization, pre-training, post-training, and inference. Strikingly, bypassing traditional subword segmentation via a character-level counterfactual tokenizer removes neither input and output asymmetries nor dialectal accuracy gaps. During pre-training, dialect pairs induce more divergent gradient updates than pairs of entirely unrelated SAE documents, indicating that models find semantically equivalent dialectal content harder to learn from than unrelated SAE documents. During post-training, reward models show contextual, unstable dialect preferences, assigning higher values to isolated AAVE-exclusive tokens than to SAE-exclusive tokens, while full reasoning contexts receive task- and model-dependent dialect penalties. Overall, our findings suggest that the dialect tax is encoded and accumulated not by any one step in isolation, but at every step of the language modeling process.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。