针对南亚文化偏见,构建多语言偏见词典并评估大模型生成中的性别与宗教偏见。
Purdah and Patriarchy: Evaluating and Mitigating South Asian Biases in Open-Ended Multilingual LLM Generations
- 构建涵盖性别、宗教、婚育等维度的文化偏见词典,覆盖10种南亚语言。
- 发现大模型在故事创作等任务中隐性强化了面纱制和父权制文化偏见。
- 验证提示工程可减轻文化偏见,尤其适用于印地语和泰米尔语等语言。
大型语言模型的评估常忽视交叉性与文化特异性偏见,尤其在南亚等多语言地区。本文对10种印度-雅利安语和德拉威语的LLM生成内容进行交叉文化分析,揭示面纱制与父权制文化偏见如何在开放式生成任务(如讲故事、兴趣爱好、待办清单)中被强化。研究构建了一个文化根基的偏见词典,涵盖性别、宗教、婚姻状况、子女数量等未被充分探索的交叉维度,并用于量化偏见及自去偏策略的有效性。进一步评估了简单与复杂提示两种自去偏方法在减少南亚语言中文化特定偏见方面的效果。该方法通过新颖的偏见词典与评估框架,为超越欧洲中心化或小规模多语言场景提供了更细致的文化偏见洞察。
原文摘要 · Abstract (English)
Evaluations of Large Language Models (LLMs) often overlook intersectional and culturally specific biases, particularly in underrepresented multilingual regions like South Asia. This work addresses these gaps by conducting a multilingual and intersectional analysis of LLM outputs across 10 Indo-Aryan and Dravidian languages, identifying how cultural stigmas influenced by purdah and patriarchy are reinforced in generative tasks. We construct a culturally grounded bias lexicon capturing previously unexplored intersectional dimensions including gender, religion, marital status, and number of children. We use our lexicon to quantify intersectional bias and the effectiveness of self-debiasing in open-ended generations (e.g., storytelling, hobbies, and to-do lists), where bias manifests subtly and remains largely unexamined in multilingual contexts. Finally, we evaluate two self-debiasing strategies (simple and complex prompts) to measure their effectiveness in reducing culturally specific bias in Indo-Aryan and Dravidian languages. Our approach offers a nuanced lens into cultural bias by introducing a novel bias lexicon and evaluation framework that extends beyond Eurocentric or small-scale multilingual settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。