arXiv:2603.15187cs.CL2026-03被引 1

研究英语方言公平性,发现数据少导致模型难提升

The Hrunting of AI: Where and How to Improve English Dialectal Fairness

  • 用人类一致性评估模型生成质量,发现一致率影响模型判断
  • 方言区人类共识低时,模型表现更差且微调可能加剧问题
  • 部分模型可生成高质量数据,为小众方言提升带来希望

大型语言模型在英语方言上的表现不佳,且因数据稀缺难以改进。本文评估了三种较少研究的英语方言(约克郡、盖尔迪、康沃尔)及非裔美国黑人英语、西弗里斯兰语作为对照。发现人类间对模型生成质量的判断一致性直接影响模型作为裁判的表现;模型与人类的一致性模式与人类间一致性一致,准确率等指标也呈现类似趋势。这说明模型表现依赖于人类共识,而人口稀少地区共识本就较低,从而限制了模型改进的可行性。微调无法消除这一现象,甚至可能放大。但观察到部分模型能生成高质量数据,具备规模化潜力。因此,需谨慎评估数据质量以实现公平包容的模型改进;在数据稀缺时,亟需新工具应对该模式。

原文摘要 · Abstract (English)

It is known that large language models (LLMs) underperform in English dialects, and that improving them is difficult due to data scarcity. In this work we investigate how quality and availability impact the feasibility of improving LLMs in this context. For this, we evaluate three rarely-studied English dialects (Yorkshire, Geordie, and Cornish), plus African-American Vernacular English, and West Frisian as control. We find that human-human agreement when determining LLM generation quality directly impacts LLM-as-a-judge performance. That is, LLM-human agreement mimics the human-human agreement pattern, and so do metrics such as accuracy. It is an issue because LLM-human agreement measures an LLM's alignment with the human consensus; and hence raises questions about the feasibility of improving LLM performance in locales where low populations induce low agreement. We also note that fine-tuning does not eradicate, and might amplify, this pattern in English dialects. But also find encouraging signals, such as some LLMs' ability to generate high-quality data, thus enabling scalability. We argue that data must be carefully evaluated to ensure fair and inclusive LLM improvement; and, in the presence of scarcity, new tools are needed to handle the pattern found.

方言公平性大模型评估数据稀缺

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。