arXiv:2505.20875cs.CLcs.AI2025-05NeurIPS被引 5

构建跨英语变体评估框架,揭示大模型在非标准英语上的性能下降46.3%。

Trans-EnV: A Framework for Evaluating the Linguistic Robustness of LLMs Against English Varieties

  • 基于语言学知识与大模型生成,自动将标准英语数据转为38种英语变体
  • 7个主流大模型在非标准变体上准确率最高下降46.3%
  • 适合关注AI公平性、多语种应用的研究者与开发者

大型语言模型(LLMs)主要在标准美式英语(SAE)上进行评估,忽视了全球英语变体的多样性。这种局限可能引发公平性问题,因非标准变体上性能下降会导致全球用户受益不均。因此,全面评估模型对多种非标准英语变体的鲁棒性至关重要。我们提出Trans-EnV框架,能自动将SAE数据集转化为多个英语变体以评估语言鲁棒性。该框架结合(1)语言学专家知识,从文献和语料库中提炼变体特异性特征与转换规则;(2)基于LLM的转换方法,兼顾语言有效性与可扩展性。利用Trans-EnV,我们将6个基准数据集转化为38种英语变体,并评估了7个最先进的LLMs。结果表明存在显著性能差异,非标准变体上准确率最高下降46.3%。每项转换均通过统计检验及二语习得领域研究者审核,确保语言有效性。代码与数据集已公开于https://github.com/jiyounglee-0523/TransEnV 和 https://huggingface.co/collections/jiyounglee0523/transenv-681eadb3c0c8cf363b363fb1。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are predominantly evaluated on Standard American English (SAE), often overlooking the diversity of global English varieties. This narrow focus may raise fairness concerns as degraded performance on non-standard varieties can lead to unequal benefits for users worldwide. Therefore, it is critical to extensively evaluate the linguistic robustness of LLMs on multiple non-standard English varieties. We introduce Trans-EnV, a framework that automatically transforms SAE datasets into multiple English varieties to evaluate the linguistic robustness. Our framework combines (1) linguistics expert knowledge to curate variety-specific features and transformation guidelines from linguistic literature and corpora, and (2) LLM-based transformations to ensure both linguistic validity and scalability. Using Trans-EnV, we transform six benchmark datasets into 38 English varieties and evaluate seven state-of-the-art LLMs. Our results reveal significant performance disparities, with accuracy decreasing by up to 46.3% on non-standard varieties. These findings highlight the importance of comprehensive linguistic robustness evaluation across diverse English varieties. Each construction of Trans-EnV was validated through rigorous statistical testing and consultation with a researcher in the field of second language acquisition, ensuring its linguistic validity. Our code and datasets are publicly available at https://github.com/jiyounglee-0523/TransEnV and https://huggingface.co/collections/jiyounglee0523/transenv-681eadb3c0c8cf363b363fb1.

语言模型英语变体公平性鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。