针对印度文化背景设计新数据集与评估框架,发现大模型在残障相关偏见上尤为严重。
IndiCASA: A Dataset and Bias Evaluation Framework in LLMs Using Contrastive Embedding Similarity in the Indian Context
- 用对比学习训练编码器,通过嵌入相似性捕捉细微偏见
- 在2575条经人工验证语句上测试,所有模型均存在刻板印象偏差
- 适合关注本地化大模型公平性的研究者和开发者
大语言模型在关键领域广泛应用,但其在文化多元的印度等地区部署时,现有基于嵌入的偏见评估方法难以捕捉细微的刻板印象。本文提出一种基于对比学习训练编码器的评估框架,可捕获细粒度偏见。同时构建了新数据集IndiCASA(印度偏见相关情境对齐刻板印象与反刻板印象),包含2575条人工验证句子,覆盖种姓、性别、宗教、残疾和经济地位五个维度。对多个开源大模型的评估显示,所有模型均存在一定程度的刻板印象偏差,其中与残障相关的偏见尤为顽固;宗教偏见普遍较低,可能得益于全球范围的去偏努力。结果凸显了在本土化场景下开发更公平模型的必要性。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have gained significant traction across critical domains owing to their impressive contextual understanding and generative capabilities. However, their increasing deployment in high stakes applications necessitates rigorous evaluation of embedded biases, particularly in culturally diverse contexts like India where existing embedding-based bias assessment methods often fall short in capturing nuanced stereotypes. We propose an evaluation framework based on a encoder trained using contrastive learning that captures fine-grained bias through embedding similarity. We also introduce a novel dataset - IndiCASA (IndiBias-based Contextually Aligned Stereotypes and Anti-stereotypes) comprising 2,575 human-validated sentences spanning five demographic axes: caste, gender, religion, disability, and socioeconomic status. Our evaluation of multiple open-weight LLMs reveals that all models exhibit some degree of stereotypical bias, with disability related biases being notably persistent, and religion bias generally lower likely due to global debiasing efforts demonstrating the need for fairer model development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。