arXiv:2512.19238cs.CLcs.AI2025-12AAAI被引 1

研究93个受污名化群体在大模型中的偏见,发现危险性高的群体偏见最严重。

Identifying Features Associated with Bias Against 93 Stigmatized Groups in Language Models and Guardrail Model Safety Mitigation

  • 基于六类社会特征分析污名化群体的偏见成因
  • 高危险性污名群体输出偏见率达60%,社会属性群体仅11%
  • 防护模型可减少偏见但无法识别意图,需改进

大型语言模型(LLMs)存在社会偏见,但对非受保护的污名化身份的研究仍不足。心理学研究表明,污名包含六种共享社会特征:审美性、隐蔽性、病程、破坏性、起源和危险性。本研究探究人类与模型对这些特征的评分、提示风格及污名类型如何影响模型输出中的偏见。通过SocialStigmaQA基准测试,评估三种主流模型(Granite 3.0-8B、Llama-3.1-8B、Mistral-7B)在93个污名化群体上的表现,涵盖37个社会情境(如是否推荐实习)。结果表明,人类评定为高危险性的污名(如黑帮成员或艾滋病患者)导致模型输出偏见率高达60%,而社会人口学特征污名(如亚裔美国人或老年人)仅11%。进一步测试各模型对应的防护模型(Granite Guardian 3.0、Llama Guard 3.0、Mistral Moderation API)能否缓解偏见,结果显示偏见分别下降10.4%、1.4%和7.8%。然而,防护后显著影响偏见的关键特征依然不变,且防护模型常无法识别提示中的偏见意图。该研究对涉及污名群体的应用有重要启示,建议未来提升防护模型对偏见意图的识别能力。

原文摘要 · Abstract (English)

Large language models (LLMs) have been shown to exhibit social bias, however, bias towards non-protected stigmatized identities remain understudied. Furthermore, what social features of stigmas are associated with bias in LLM outputs is unknown. From psychology literature, it has been shown that stigmas contain six shared social features: aesthetics, concealability, course, disruptiveness, origin, and peril. In this study, we investigate if human and LLM ratings of the features of stigmas, along with prompt style and type of stigma, have effect on bias towards stigmatized groups in LLM outputs. We measure bias against 93 stigmatized groups across three widely used LLMs (Granite 3.0-8B, Llama-3.1-8B, Mistral-7B) using SocialStigmaQA, a benchmark that includes 37 social scenarios about stigmatized identities; for example deciding wether to recommend them for an internship. We find that stigmas rated by humans to be highly perilous (e.g., being a gang member or having HIV) have the most biased outputs from SocialStigmaQA prompts (60% of outputs from all models) while sociodemographic stigmas (e.g. Asian-American or old age) have the least amount of biased outputs (11%). We test if the amount of biased outputs could be decreased by using guardrail models, models meant to identify harmful input, using each LLM's respective guardrail model (Granite Guardian 3.0, Llama Guard 3.0, Mistral Moderation API). We find that bias decreases significantly by 10.4%, 1.4%, and 7.8%, respectively. However, we show that features with significant effect on bias remain unchanged post-mitigation and that guardrail models often fail to recognize the intent of bias in prompts. This work has implications for using LLMs in scenarios involving stigmatized groups and we suggest future work towards improving guardrail models for bias mitigation.

语言模型偏见检测防护机制社会偏见

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。