arXiv:2507.19195cs.CLcs.AI2025-07

小规模数据污染可放大方言偏见,尤其影响非裔英语使用者

Can Small-Scale Data Poisoning Exacerbate Dialect-Linked Biases in Large Language Models?

  • 用少量方言提示+负面回复污染训练数据,诱导模型产生偏见
  • 非裔英语输入时毒性与刻板印象显著上升,标准英语也受影响
  • 传统检测易漏判,需关注内容层面的刻板印象风险

风格条件式数据污染被识别为放大大型语言模型社会语言偏见的隐蔽途径。通过在指令微调中使用少量污染数据,将非裔美国黑人俚语(AAVE)和南方方言提示与有毒或刻板回应配对,本研究探究语言风格是否可作为有害行为的潜在触发因素。在多个模型家族与规模下,污染暴露导致方言输入的毒性与刻板印象表达显著升高,尤以AAVE最为明显;标准美式英语虽相对较低,但并非免疫。多指标审计结合分类器毒性检测与大模型作为裁判的方法揭示,即使词汇毒性不明显,仍存在大量刻板印象内容,表明传统检测低估了社会语言危害。此外,污染模型表现出未明确包含辱骂词的新兴越狱行为,暗示对齐能力下降而非单纯记忆。研究强调需开展方言敏感评估、内容级刻板印象审计,以及显式解耦风格与毒性的训练协议,以防范看似微小的风格污染引发的偏见放大。

原文摘要 · Abstract (English)

Style-conditioned data poisoning is identified as a covert vector for amplifying sociolinguistic bias in large language models. Using small poisoned budgets that pair dialectal prompts -- principally African American Vernacular English (AAVE) and a Southern dialect -- with toxic or stereotyped completions during instruction tuning, this work probes whether linguistic style can act as a latent trigger for harmful behavior. Across multiple model families and scales, poisoned exposure elevates toxicity and stereotype expression for dialectal inputs -- most consistently for AAVE -- while Standard American English remains comparatively lower yet not immune. A multi-metric audit combining classifier-based toxicity with an LLM-as-a-judge reveals stereotype-laden content even when lexical toxicity appears muted, indicating that conventional detectors under-estimate sociolinguistic harms. Additionally, poisoned models exhibit emergent jailbreaking despite the absence of explicit slurs in the poison, suggesting weakened alignment rather than memorization. These findings underscore the need for dialect-aware evaluation, content-level stereotype auditing, and training protocols that explicitly decouple style from toxicity to prevent bias amplification through seemingly minor, style-based contamination.

语言模型偏见放大方言数据污染

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。