arXiv:2510.06974cs.CL2025-10ACL

中文大模型存在群体身份偏见,性别标记代词加剧毒性

Probing Social Identity Bias in Chinese LLMs with Gendered Pronouns and Social Groups

  • 用汉语特有代词设计对比实验,区分默认中性与女性标记复数代词
  • 10个中文模型中均发现群体内偏爱现象,女性标记代词毒性更高
  • 首次揭示汉语语法结构可暴露英文模型难以捕捉的偏见模式

大型语言模型在用户应用中日益普及,引发对其可能反映并放大社会偏见的担忧。本文针对十款代表性中文大模型,使用普通话特有的提示语,考察240个在中国语境下显著的社会群体在‘我们’(ingroup)与‘他们’(outgroup)框架下的偏见表现。评估采用双层测量框架,分别衡量情感倾向与毒性水平。提示设计充分考虑汉语语言特性,尤其是默认中性复数代词与其显性女性对应形式之间的区别,从而实现对社会身份框架效应的受控比较。结果显示,所有模型均存在系统性的群体内-外不对称现象,但不同评估维度表现各异:指令微调常能缓解情感偏差,而毒性差距则更为顽固。此外,在多个模型中,女性标记的复数代词比默认中性复数代词关联更高的毒性。本研究提出面向中文大模型的语言感知评估框架,表明(i)英语中已知的社会身份偏见同样存在于中文场景;(ii)汉语特有的语言结构可揭示英文语境下无法直接观察到的偏见模式。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly deployed in user-facing applications, raising concerns that they may reflect and amplify social biases. We investigate social identity biases in Chinese LLMs using Mandarin-specific prompts across ten representative models. Our evaluation compares ingroup ("We") and outgroup ("They") framings across 240 social groups salient in the Chinese context, using a two-tiered measurement framework that assesses both sentiment and toxicity. The prompt design explicitly accounts for linguistic properties of Mandarin, including the distinction between the default gender-neutral plural pronoun and its explicitly feminine counterpart, enabling a controlled comparison of social identity framing effects. Across models, we observe systematic ingroup-outgroup asymmetries, although their expression differs across measurement dimensions. In particular, instruction tuning often reduces sentiment asymmetries, while toxicity gaps remain more persistent. Moreover, the feminine-marked plural pronoun is associated with higher toxicity than the default gender-neutral plural in several models. Our study introduces a language-aware evaluation framework for Chinese LLMs and shows that (i) social identity biases previously documented in English also manifest in Chinese and that (ii) Mandarin-specific linguistic structure can reveal bias patterns that are not directly observable in English-only settings.

中文大模型社会偏见语言结构毒性检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。