在尼泊尔文化背景下,发现大模型存在显性与隐性偏见,且二者表现不一致。
Dual-Metric Evaluation of Social Bias in Large Language Models: Evidence from an Underrepresented Nepali Cultural Context
- 设计双指标评估框架,同时检测模型对刻板印象的认同和生成倾向。
- 模型显性偏见率0.36-0.43,隐性生成偏见率达0.74-0.755,且随温度变化呈非线性波动。
- 隐性偏见更难被传统评估捕捉,适合关注文化偏见与模型公平性的研究者。
大型语言模型(LLMs)日益影响全球数字生态,但其在欠代表文化语境中可能延续社会与文化偏见的问题仍不清楚。本研究系统分析了七款先进大模型(GPT-4o-mini、Claude-3-Sonnet、Claude-4-Sonnet、Gemini-2.0-Flash、Gemini-2.0-Lite、Llama-3-70B、Mistral-Nemo)在尼泊尔文化背景下的表征偏见。基于超过2400个关于性别角色的社会领域刻板与反刻板语句对,构建了符合Croissant标准的数据集,并提出双指标偏见评估框架(DMBA),结合两项指标:(1)对偏见陈述的认同度,(2)刻板化补全倾向。结果显示,模型显性认同偏见的均值为0.36–0.43,隐性生成偏见率为0.740–0.755。值得注意的是,隐性偏见与温度呈非线性U型关系,在中等随机性(T=0.3)时达到峰值,高温下略有下降。不同解码设置下的相关性分析表明,显性认同与刻板句认同高度相关,但对隐性生成偏见的预测能力弱且常为负相关,说明生成偏见无法被认同度有效捕获。敏感性分析显示,增加top-p会放大显性偏见,而隐性生成偏见基本保持稳定。领域层面分析发现,隐性偏见在种族与社会文化刻板印象中最强,显性认同偏见在性别与社会文化类别间相似,种族类别的显性认同最低。研究强调需为欠代表社会构建文化适配数据集与去偏策略。
原文摘要 · Abstract (English)
Large language models (LLMs) increasingly influence global digital ecosystems, yet their potential to perpetuate social and cultural biases remains poorly understood in underrepresented contexts. This study presents a systematic analysis of representational biases in seven state-of-the-art LLMs: GPT-4o-mini, Claude-3-Sonnet, Claude-4-Sonnet, Gemini-2.0-Flash, Gemini-2.0-Lite, Llama-3-70B, and Mistral-Nemo in the Nepali cultural context. Using Croissant-compliant dataset of 2400+ stereotypical and anti-stereotypical sentence pairs on gender roles across social domains, we implement an evaluation framework, Dual-Metric Bias Assessment (DMBA), combining two metrics: (1) agreement with biased statements and (2) stereotypical completion tendencies. Results show models exhibit measurable explicit agreement bias, with mean bias agreement ranging from 0.36 to 0.43 across decoding configurations, and an implicit completion bias rate of 0.740-0.755. Importantly, implicit completion bias follows a non-linear, U-shaped relationship with temperature, peaking at moderate stochasticity (T=0.3) and declining slightly at higher temperatures. Correlation analysis under different decoding settings revealed that explicit agreement strongly aligns with stereotypical sentence agreement but is a weak and often negative predictor of implicit completion bias, indicating generative bias is poorly captured by agreement metrics. Sensitivity analysis shows increasing top-p amplifies explicit bias, while implicit generative bias remains largely stable. Domain-level analysis shows implicit bias is strongest for race and sociocultural stereotypes, while explicit agreement bias is similar across gender and sociocultural categories, with race showing the lowest explicit agreement. These findings highlight the need for culturally grounded datasets and debiasing strategies for LLMs in underrepresented societies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。