arXiv:2504.07100cs.CL2025-04EMNLP被引 5

构建英语方言评测集,揭示大模型在非标准语境下的性能短板

EnDive: A Cross-Dialect Benchmark for Fairness and Performance in Large Language Models

  • 用少量示例将标准英语数据转译为5种弱势方言
  • 模型在方言输入上表现显著低于标准英语,差距超15%
  • 适合关注语言公平性与模型鲁棒性的研究者使用

人类语言的多样性受社会、文化与地域影响,给自然语言处理系统带来挑战。现有评测基准常忽视语言内部差异,导致非标准方言使用者被边缘化。为此,我们提出EnDive(英语多样性)基准,评估五种主流大语言模型在语言理解、算法推理、数学和逻辑任务中的表现。通过少样本提示并结合母语者验证示例,将标准美式英语数据集翻译成五种代表性弱势方言,并以流畅度、偏好测试和语义相似性指标对比规则方法。人工评估显示翻译质量高,忠实度、流畅度与正式度平均得分均达6.02/7以上。剔除近似重复后,构建出具有挑战性的数据集,揭示模型在方言输入上持续表现不佳,性能差距超过15%。EnDive推动了更具方言意识的NLP发展,暴露模型偏见,促进更公平的语言技术。

原文摘要 · Abstract (English)

The diversity of human language, shaped by social, cultural, and regional influences, presents significant challenges for natural language processing (NLP) systems. Existing benchmarks often overlook intra-language variations, leaving speakers of non-standard dialects underserved. To address this gap, we introduce EnDive (English Diversity), a benchmark that evaluates five widely-used large language models (LLMs) across tasks in language understanding, algorithmic reasoning, mathematics, and logic. Our framework translates Standard American English datasets into five underrepresented dialects using few-shot prompting with verified examples from native speakers, and compare these translations against rule-based methods via fluency assessments, preference tests, and semantic similarity metrics. Human evaluations confirm high translation quality, with average scores of at least 6.02/7 for faithfulness, fluency, and formality. By filtering out near-identical translations, we create a challenging dataset that reveals significant performance disparities - models consistently underperform on dialectal inputs compared to Standard American English. EnDive thus advances dialect-aware NLP by uncovering model biases and promoting more equitable language technologies.

语言公平大模型评测方言识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。