检测多语言大模型在波斯语中的性别偏见,发现其偏见比英语更严重。
Probing Gender Bias in Multilingual LLMs: A Case Study of Stereotypes in Persian
- 设计模板探测法结合领域特定性别偏差指数,量化模型偏见程度。
- 四模型在波斯语中均存在性别刻板印象,体育领域偏差最明显。
- 为低资源语言的公平性评估提供可复用框架,适合关注AI伦理的研究者。
多语言大语言模型在全球广泛应用,确保其无性别偏见对避免表征伤害至关重要。尽管已有研究关注高资源语言的偏见问题,但低资源语言仍缺乏深入探讨。本文提出一种基于模板的探测方法,并通过真实数据验证,用于揭示多语言模型中的性别刻板印象。为此,我们引入领域特定性别偏差指数(DS-GSI),量化性别平等偏离程度。我们在四个语义领域评估了GPT-4o mini、DeepSeek R1、Gemini 2.0 Flash和Qwen QwQ 32B四款主流模型,聚焦波斯语——一种具有独特语言特征的低资源语言。结果表明,所有模型均存在性别刻板印象,且在波斯语中的偏差大于英语,其中体育领域表现最为僵化。本研究强调了包容性NLP实践的必要性,并提供了评估其他低资源语言偏见的可推广框架。
原文摘要 · Abstract (English)
Multilingual Large Language Models (LLMs) are increasingly used worldwide, making it essential to ensure they are free from gender bias to prevent representational harm. While prior studies have examined such biases in high-resource languages, low-resource languages remain understudied. In this paper, we propose a template-based probing methodology, validated against real-world data, to uncover gender stereotypes in LLMs. As part of this framework, we introduce the Domain-Specific Gender Skew Index (DS-GSI), a metric that quantifies deviations from gender parity. We evaluate four prominent models, GPT-4o mini, DeepSeek R1, Gemini 2.0 Flash, and Qwen QwQ 32B, across four semantic domains, focusing on Persian, a low-resource language with distinct linguistic features. Our results show that all models exhibit gender stereotypes, with greater disparities in Persian than in English across all domains. Among these, sports reflect the most rigid gender biases. This study underscores the need for inclusive NLP practices and provides a framework for assessing bias in other low-resource languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。