首个系统评估大模型在方言推理任务中公平性的研究,发现主流模型对非标准语体严重不公。
Assessing Dialect Fairness and Robustness of Large Language Models in Reasoning Tasks
- 构建平行语料库ReDial,包含1200+组标准英语与非洲裔美国黑人英语对照查询
- 多数大模型在AAVE queries上表现显著下降,推理准确率普遍低于标准语体
- 面向语言公平性研究者与开发者,推动更包容的AI系统设计
语言并非单一整体。尽管已有多种多语言基准用于评估大语言模型(LLMs)性能,但它们往往忽视同一语言内部的差异,难以反映非标准方言使用者的真实体验。本文聚焦非洲裔美国黑人英语(AAVE),首次系统评估主流大模型在算法、数学、逻辑及综合推理等典型任务中对方言的公平性与鲁棒性。我们构建了ReDial(Reasoning with Dialect Queries)基准,包含1200+组标准英语与AAVE的平行查询对。通过雇佣具备计算机科学背景的AAVE母语者,将HumanEval、GSM8K等七项流行基准改写为AAVE版本。评估涵盖GPT、Claude、Llama、Mistral及Phi系列等主流模型。结果表明,几乎所有模型在处理AAVE查询时均表现出显著脆弱性和不公平性。本研究建立了分析大模型方言偏见的系统化客观框架,揭示了主流模型在推理任务中对方言使用者的不公服务,为未来研究奠定关键基础。
原文摘要 · Abstract (English)
Language is not monolithic. While benchmarks, including those designed for multiple languages, are often used as proxies to evaluate the performance of Large Language Models (LLMs), they tend to overlook the nuances of within-language variation and thus fail to model the experience of speakers of non-standard dialects. Focusing on African American Vernacular English (AAVE), we present the first study aimed at objectively assessing the fairness and robustness of LLMs in handling dialects across canonical reasoning tasks, including algorithm, math, logic, and integrated reasoning. We introduce ReDial (Reasoning with Dialect Queries), a benchmark containing 1.2K+ parallel query pairs in Standardized English and AAVE. We hire AAVE speakers, including experts with computer science backgrounds, to rewrite seven popular benchmarks, such as HumanEval and GSM8K. With ReDial, we evaluate widely used LLMs, including GPT, Claude, Llama, Mistral, and the Phi model families. Our findings reveal that almost all of these widely used models show significant brittleness and unfairness to queries in AAVE. Our work establishes a systematic and objective framework for analyzing LLM bias in dialectal queries. Moreover, it highlights how mainstream LLMs provide unfair service to dialect speakers in reasoning tasks, laying a critical foundation for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。