跨语言评估大模型偏见,发现阿拉伯语和西班牙语偏见更高
Cross-Language Bias Examination in Large Language Models
- 结合显性测试与隐性测验,多语言对比偏见程度
- 阿拉伯语/西班牙语显性偏见低但隐性偏见高,年龄相关隐性偏见最严重
- 为构建公平多语言大模型提供可复现的评估框架
本研究提出一种创新的多语言偏见评估框架,结合显性偏见评估(通过BBQ基准)与基于提示的隐性关联测试(prompt-based IAT),将提示和词表翻译成英语、中文、阿拉伯语、法语和西班牙语,直接比较不同语言中的偏见差异。结果显示,大模型在不同语言间存在显著偏见差距:阿拉伯语和西班牙语使用者表现出更高的刻板印象偏见,而中文和英语相对较低。同时发现偏见类型模式不一:年龄相关的显性偏见最低,但隐性偏见最高,凸显标准基准无法捕捉隐性偏见的重要性。研究揭示大模型在语言和偏见维度上差异显著,填补了跨语言偏见分析的研究空白,为开发公平多语言大模型奠定基础。
原文摘要 · Abstract (English)
This study introduces an innovative multilingual bias evaluation framework for assessing bias in Large Language Models, combining explicit bias assessment through the BBQ benchmark with implicit bias measurement using a prompt-based Implicit Association Test. By translating the prompts and word list into five target languages, English, Chinese, Arabic, French, and Spanish, we directly compare different types of bias across languages. The results reveal substantial gaps in bias across languages used in LLMs. For example, Arabic and Spanish consistently show higher levels of stereotype bias, while Chinese and English exhibit lower levels of bias. We also identify contrasting patterns across bias types. Age shows the lowest explicit bias but the highest implicit bias, emphasizing the importance of detecting implicit biases that are undetectable with standard benchmarks. These findings indicate that LLMs vary significantly across languages and bias dimensions. This study fills a key research gap by providing a comprehensive methodology for cross-lingual bias analysis. Ultimately, our work establishes a foundation for the development of equitable multilingual LLMs, ensuring fairness and effectiveness across diverse languages and cultures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。