arXiv:2607.06845cs.CL2026-07

LLMs常无意识纠正非裔英语,该研究提出方法检测并缓解这种偏见。

LLMs Silently Correct African American English: Auditing and Mitigating Dialect Bias via Activation Steering

论文配图:LLMs Silently Correct African American English: Auditing and Mitigating Dialect Bias via Activation Steering
图 1 · 摘自论文原文
  • 用条件方言不变性检测模型真实偏见,排除翻译干扰
  • 发现否定一致等句法特征是触发偏见的核心因素
  • 首次应用激活转向技术,比提示词更有效且不损害标准英语流畅性

非裔美国人英语(AAE)是超过3000万人使用的规则化方言,却常被大型语言模型错误解读并“修正”为标准美式英语(SAE)。在六种指令微调的LLM(14B至70B)中,我们发现先进模型即使上下文为AAE,仍系统性偏好SAE延续,实质上将AAE重写为SAE。为此,我们提出端到端框架用于审计与缓解偏见:审计方面,引入条件方言群不变性(cDGI),分离真实偏见与翻译伪影,并通过特征级定位分析识别出最强烈触发偏见的AAE标记,发现否定一致(如"ain't nobody")是所有模型的普遍触发点;缓解方面,首次将激活转向应用于方言偏见——一种无需训练、测试时可用的方法,通过因果追踪提取方言方向并注入偏见相关层,使偏见降低5至20倍,同时保持SAE流畅性。为支持本研究,我们发布REAL-AAE,目前最大的真实AAE平行语料库:包含17,479个自然推文的AAE/SAE/AAE_back三元组,规模达先前真实资源的2至6倍,经自动验证(BERTScore F1=0.95)和三位母语者评估(语义一致率83.0%)。

原文摘要 · Abstract (English)

African American English (AAE), a rule-governed dialect spoken by over 30 million people, is routinely misinterpreted and "corrected" by large language models (LLMs). Across six instruction-tuned LLMs (14B to 70B), we show that state-of-the-art models systematically prefer Standard American English (SAE) continuations even when the preceding context is in AAE, effectively rewriting AAE into SAE. We present an end-to-end framework to audit and mitigate this bias. For auditing, we introduce conditional Dialect Group Invariance (cDGI), which isolates true model bias from translator-induced artifacts, and a feature-level localization analysis that identifies which AAE markers most strongly trigger bias; we find that syntactic constructions, especially negative concord (e.g., "ain't nobody"), are universal triggers across all models. For mitigation, we introduce, to our knowledge, the first application of activation steering to dialect bias: a training-free, test-time method that extracts dialect directions via causal tracing and injects them into bias-relevant layers. Activation steering reduces bias 5 to 20 times more than prompting while preserving SAE fluency. To enable this work, we release REAL-AAE , the largest real-AAE parallel corpus to date: 17,479 AAE/SAE/ AAE_back triplets from natural tweets (2 to 6 times larger than prior real-AAE resources), validated automatically (BERTScore F1 = 0.95) and by three native AAE speakers (83.0% semantic agreement).

方言偏见激活转向LLM审计AAE

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。