arXiv:2603.18729cs.AI2026-03

分析大模型对美式英语与非裔英语的刻板印象差异及缓解策略

Analysis Of Linguistic Stereotypes in Single and Multi-Agent Generative AI Architectures

  • 设计八类提示模板,检测不同方言下生成内容的刻板倾向
  • 非裔英语输入常被赋予负面职业和形容词,差距最大在Claude Haiku中
  • 多代理架构比单一提示更有效,适合高风险应用部署

现有研究显示,大语言模型在输出中存在歧视性行为,会基于输入语言的方言产生刻板印象。本文复现了方言敏感性刻板印象生成的分析,并探究了提示工程(角色设定与思维链)和多代理架构(生成-批判-修正)等缓解策略的效果。设计八种提示模板,评估不同方言下生成的名字、职业、形容词等差异。采用大模型作为评判者评估偏差。结果表明,所有模板类别中,标准美式英语与非裔美国英语生成内容均存在刻板差异,尤其在形容词和职业分配上最显著。基线差异因模型而异,最大的方言差出现在Claude Haiku,最小在Phi-4 Mini。思维链提示对Claude Haiku有明显缓解作用,而多代理架构则在所有模型中保持一致的缓解效果。研究建议,在涉及交叉性公平的软件工程中,需针对不同模型验证缓解策略,并在高影响场景中引入工作流级控制(如包含批判模型的智能体架构)。当前结果为探索性,但可扩展至更多语言或方言。

原文摘要 · Abstract (English)

Many works in the literature show that LLM outputs exhibit discriminatory behaviour, triggering stereotype-based inferences based on the dialect in which the inputs are written. This bias has been shown to be particularly pronounced when the same inputs are provided to LLMs in Standard American English (SAE) and African-American English (AAE). In this paper, we replicate existing analyses of dialect-sensitive stereotype generation in LLM outputs and investigate the effects of mitigation strategies, including prompt engineering (role-based and Chain-Of-Thought prompting) and multi-agent architectures composed of generate-critique-revise models. We define eight prompt templates to analyse different ways in which dialect bias can manifest, such as suggested names, jobs, and adjectives for SAE or AAE speakers. We use an LLM-as-judge approach to evaluate the bias in the results. Our results show that stereotype-bearing differences emerge between SAE- and AAE-related outputs across all template categories, with the strongest effects observed in adjective and job attribution. Baseline disparities vary substantially by model, with the largest SAE-AAE differential observed in Claude Haiku and the smallest in Phi-4 Mini. Chain-Of-Thought prompting proved to be an effective mitigation strategy for Claude Haiku, whereas the use of a multi-agent architecture ensured consistent mitigation across all the models. These findings suggest that for intersectionality-informed software engineering, fairness evaluation should include model-specific validation of mitigation strategies, and workflow-level controls (e.g., agentic architectures involving critique models) in high-impact LLM deployments. The current results are exploratory in nature and limited in scope, but can lead to extensions and replications by increasing the dataset size and applying the procedure to different languages or dialects.

语言偏见大模型多代理公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。