arXiv:2506.23111cs.CL2025-06ACL被引 8

针对印度多元文化背景,构建首个本土化大模型公平性评估基准。

FairI Tales: Evaluation of Fairness in Indian Contexts with a Focus on Bias and Stereotypes

  • 基于1800个社会文化话题生成2万条真实场景模板,覆盖85类身份群体。
  • 14款主流大模型在测试中普遍对边缘群体存在偏见,且难以自我纠正。
  • 适用于关注算法公平性、跨文化AI研究及政策制定者。

现有公平性研究多聚焦西方语境,难以适用于印度等文化多样性国家。为此,我们提出INDIC-BIAS,一个面向印度的综合性基准,用于评估大模型在85类身份群体(涵盖种姓、宗教、地区、部落)上的公平性。首先,通过领域专家协作,整理超过1800个可能滋生偏见与刻板印象的社会文化主题。在此基础上,生成并人工验证了2万条真实世界情境模板,用于探测大模型的公平性表现。这些模板被结构化为三类任务:合理性判断、价值判断与内容生成。对14款主流大模型的评估显示,模型普遍对边缘身份群体存在显著负面偏见,且在被要求解释决策时仍难以缓解偏见。结果揭示了当前大模型可能引发的分配性与表征性伤害,警示其在实际应用中需更加审慎。我们开源INDIC-BIAS,以推动印度语境下偏见与刻板印象的评测与缓解研究。

原文摘要 · Abstract (English)

Existing studies on fairness are largely Western-focused, making them inadequate for culturally diverse countries such as India. To address this gap, we introduce INDIC-BIAS, a comprehensive India-centric benchmark designed to evaluate fairness of LLMs across 85 identity groups encompassing diverse castes, religions, regions, and tribes. We first consult domain experts to curate over 1,800 socio-cultural topics spanning behaviors and situations, where biases and stereotypes are likely to emerge. Grounded in these topics, we generate and manually validate 20,000 real-world scenario templates to probe LLMs for fairness. We structure these templates into three evaluation tasks: plausibility, judgment, and generation. Our evaluation of 14 popular LLMs on these tasks reveals strong negative biases against marginalized identities, with models frequently reinforcing common stereotypes. Additionally, we find that models struggle to mitigate bias even when explicitly asked to rationalize their decision. Our evaluation provides evidence of both allocative and representational harms that current LLMs could cause towards Indian identities, calling for a more cautious usage in practical applications. We release INDIC-BIAS as an open-source benchmark to advance research on benchmarking and mitigating biases and stereotypes in the Indian context.

大模型公平性印度语境偏见评估社会影响

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。