arXiv:2503.04099cs.CLcs.AI2025-03被引 10

研究发现大模型对非裔英语的推理准确率更低,解释更简单。

Disparities in LLM Reasoning Accuracy and Explanations: A Case Study on African American English

  • 用标准英语和非裔英语对比测试模型推理能力
  • 非裔英语输入时准确率下降,推理链条更短
  • 社会与人文学科领域差异最明显,适合关注公平性的研究者

大型语言模型(LLMs)在推理任务中表现出色,但近期研究揭示其在处理方言如非裔美国人英语(AAE)时存在显著偏见。本文系统研究了模型在不同语言变体下的推理差异,构建实验框架,通过基于LLM的方言转换结合已有语言学分析,比较标准美式英语(SAE)与AAE提示下的模型表现。结果表明,对于等效的非裔英语输入,模型产生的回答准确率更低,推理链更简短,解释更薄弱,尤其在社会科学和人文学科领域差异最为突出。这揭示出模型对不同语言变体处理方式存在系统性差异,引发关于多语言多方言世界中模型开发与部署的深层反思。代码已公开于 https://github.com/Runtaozhou/dialect_bias_eval。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated remarkable capabilities in reasoning tasks, leading to their widespread deployment. However, recent studies have highlighted concerning biases in these models, particularly in their handling of dialectal variations like African American English (AAE). In this work, we systematically investigate dialectal disparities in LLM reasoning tasks. We develop an experimental framework comparing LLM performance given Standard American English (SAE) and AAE prompts, combining LLM-based dialect conversion with established linguistic analyses. We find that LLMs consistently produce less accurate responses and simpler reasoning chains and explanations for AAE inputs compared to equivalent SAE questions, with disparities most pronounced in social science and humanities domains. These findings highlight systematic differences in how LLMs process and reason about different language varieties, raising important questions about the development and deployment of these systems in our multilingual and multidialectal world. Our code repository is publicly available at https://github.com/Runtaozhou/dialect_bias_eval.

大模型偏见非裔英语推理差距语言多样性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。