arXiv:2510.00962cs.CL2025-10EMNLP被引 5

研究大模型在方言上的表现偏差,发现标准英语问题改写后准确率下降20%。

Analyzing Dialectical Biases in LLMs for Knowledge and Reasoning Benchmarks

  • 将标准英语问题改写为非标准方言形式测试模型表现
  • 三种语法结构导致准确率最高下降20%:存在句'it'、零系动词、y'all
  • 提示未来应针对关键语法结构设计去偏方法,适合语言公平性研究者

大型语言模型(LLMs)在现代自然语言处理中广泛应用。然而,已有研究表明,模型在非主流英语方言上的表现会下降。本文分析了将标准美式英语问题转换为非标准方言变体对多项选择题问答任务的影响,发现准确率最高下降20%。同时,我们探究了非标准英语问题中表现不佳的语法成因。结果表明,个别语法规则对性能影响不同,其中三个特定规则——存在句中的'it'、零系动词(zero copula)以及'y'all'——能解释多数方言下性能下降现象。我们呼吁未来工作聚焦于针对高影响语法结构的偏见缓解方法。

原文摘要 · Abstract (English)

Large language models (LLMs) are ubiquitous in modern day natural language processing. However, previous work has shown degraded LLM performance for under-represented English dialects. We analyze the effects of typifying "standard" American English language questions as non-"standard" dialectal variants on multiple choice question answering tasks and find up to a 20% reduction in accuracy. Additionally, we investigate the grammatical basis of under-performance in non-"standard" English questions. We find that individual grammatical rules have varied effects on performance, but some are more consequential than others: three specific grammar rules (existential "it", zero copula, and y'all) can explain the majority of performance degradation observed in multiple dialects. We call for future work to investigate bias mitigation methods focused on individual, high-impact grammatical structures.

语言模型方言偏差语法影响公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。