揭示大模型中的种姓偏见,评估其在社会、经济等多维度的不公平表现。
DECASTE: Unveiling Caste Stereotypes in Large Language Models through Multi-Dimensional Bias Analysis
- 构建四维框架,通过定制提示检测模型对种姓群体的隐性与显性偏见。
- 实测显示达利特和首陀罗群体在模型输出中显著受歧视,偏见得分更高。
- 为关注社会公平的AI开发者提供可复用的偏见评估工具,适合伦理研究者使用。
大型语言模型(LLMs)虽在自然语言处理领域取得突破,但其仍会反映并强化社会偏见,如族裔、性别与宗教偏见。一个关键而未被充分研究的问题是种姓偏见的强化,尤其针对印度的边缘化种姓群体(如达利特人、首陀罗人)。本文提出DECASTE——一种多维度偏见分析框架,用于检测和评估LLMs中的显性和隐性种姓偏见。该框架从社会文化、经济、教育和政治四个维度出发,采用定制化提示策略进行评估。在多个前沿大模型上的基准测试表明,模型系统性地强化种姓偏见,对压迫性种姓群体的处理明显劣于主导种姓群体。例如,达利特人与首陀罗人相比,偏见评分显著更高,反映出社会偏见在模型输出中的持续存在。这些结果揭示了大模型中微妙却普遍存在的种姓偏见,强调了建立更全面、包容的偏见评估方法的必要性,以应对实际应用中的潜在风险。
原文摘要 · Abstract (English)
Recent advancements in large language models (LLMs) have revolutionized natural language processing (NLP) and expanded their applications across diverse domains. However, despite their impressive capabilities, LLMs have been shown to reflect and perpetuate harmful societal biases, including those based on ethnicity, gender, and religion. A critical and underexplored issue is the reinforcement of caste-based biases, particularly towards India's marginalized caste groups such as Dalits and Shudras. In this paper, we address this gap by proposing DECASTE, a novel, multi-dimensional framework designed to detect and assess both implicit and explicit caste biases in LLMs. Our approach evaluates caste fairness across four dimensions: socio-cultural, economic, educational, and political, using a range of customized prompting strategies. By benchmarking several state-of-the-art LLMs, we reveal that these models systematically reinforce caste biases, with significant disparities observed in the treatment of oppressed versus dominant caste groups. For example, bias scores are notably elevated when comparing Dalits and Shudras with dominant caste groups, reflecting societal prejudices that persist in model outputs. These results expose the subtle yet pervasive caste biases in LLMs and emphasize the need for more comprehensive and inclusive bias evaluation methodologies that assess the potential risks of deploying such models in real-world contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。