新基准测试大模型隐性偏见,发现种姓偏见最严重
ImplicitBBQ: Benchmarking Implicit Bias in Large Language Models through Characteristic Based Cues

- 用特征线索替代姓名,评测年龄性别等维度的隐性偏见
- 隐性偏见是显性偏见的6倍以上,种姓偏差最高达其他维度4倍
- 现有对齐和提示策略无法有效缓解隐性偏见,尤其种姓
大型语言模型在身份明确时减少偏见输出,但在身份间接表达时仍存隐性偏见。现有基准依赖姓名代理,关联弱且难以覆盖年龄、社会经济地位等维度。我们提出ImplicitBBQ,一种基于特征线索的问答基准,评估年龄、性别、地区、宗教、种姓和社会经济地位等维度的隐性偏见。评估11个模型发现,在模糊情境下隐性偏见比显性偏见高六倍以上;种姓偏见最严重,性别影响最小。安全提示和思维链推理未能显著缩小差距;即使少样本提示可降低79%隐性偏见,种姓偏差仍为其他维度的四倍。结果表明当前对齐与提示策略仅解决表层问题,深层刻板印象仍待解决。代码与数据集已公开,供研究者评估缓解方法。
原文摘要 · Abstract (English)
Large Language Models increasingly suppress biased outputs when demographic identity is stated explicitly, yet may still exhibit implicit biases when identity is conveyed indirectly. Existing benchmarks use name based proxies to detect implicit biases, which carry weak associations with many social demographics and cannot extend to dimensions like age or socioeconomic status. We introduce ImplicitBBQ, a QA benchmark that evaluates implicit bias through characteristic based cues, demographically associated attributes that signal implicitly, across age, gender, region, religion, caste, and socioeconomic status. Evaluating 11 models, we find that implicit bias in ambiguous contexts is over six times higher than explicit bias in open weight models. Notably, this bias is distributed unevenly across demographics: caste emerges as the most severe while gender is the least affected. Safety prompting and chain-of-thought reasoning fail to substantially close this gap; even few-shot prompting, which reduces implicit bias by 79%, leaves caste bias at four times the level of any other dimension. These findings indicate that current alignment and prompting strategies address the surface of bias evaluation while leaving demographically associated stereotypic associations largely unresolved. We publicly release our code and dataset for model providers and researchers to benchmark potential mitigation techniques.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。