多语言框架揭示大模型生成中的隐蔽偏见,发现所有模型均存在有害刻板印象。
StereoTales: A Multilingual Framework for Open-Ended Stereotype Discovery in LLMs

- 构建覆盖10语言、79类属性的多语言故事数据集,系统检测生成中的刻板印象
- 发现超过1500个被过度关联的刻板印象,且跨模型高度一致
- 提示语语言显著影响偏见类型,本地敏感群体更易被放大
当前多语言大模型社会偏见研究仍受限:多数基准为英语中心、模板驱动或仅限预设刻板印象识别。本文提出StereoTales,一个用于系统研究开放生成中社会偏见涌现的多语言数据集与评估流程。数据集涵盖10种语言、79个社会人口属性,包含23个近期大模型生成的65万余条故事,每条故事均标注主角19维社会人口特征。通过统计检验,识别出1500多个被过度关联的刻板印象,并由247名人类评委及同源模型评估其危害性。主要发现:(i) 所有测试模型在开放生成中均产生严重有害刻板印象,且跨厂商高度一致;(ii) 提示语语言显著影响偏见表现形式,偏见随语境文化适配,对本地敏感群体的偏见加剧;(iii) 人类与模型危害性判断总体一致(斯皮尔曼相关系数ρ=0.62),分歧集中于特定属性类别而非模型提供商。为支持后续研究,本文公开评估代码、数据集及全部标注结果。
原文摘要 · Abstract (English)
Multilingual studies of social bias in open-ended LLM generation remain limited: most existing benchmarks are English-centric, template-based, or restricted to recognizing pre-specified stereotypes. We introduce StereoTales, a multilingual dataset and evaluation pipeline for systematically studying the emergence of social bias in open-ended LLM generation. The dataset covers 10 languages and 79 socio-demographic attributes, and comprises over 650k stories generated by 23 recent LLMs, each annotated with the socio-demographic profile of the protagonist across 19 dimensions. From these, we apply statistical tests to identify more than 1{,}500 over-represented associations, which we then rate for harmfulness through both a panel of humans (N = 247) and the same LLMs. We report three main findings. \textbf{(i)} Every model we evaluate emits consequential harmful stereotypes in open-ended generation, regardless of size or capabilities, and these associations are largely shared across providers rather than isolated misbehaviors. \textbf{(ii)} Prompt language strongly shapes which stereotypes appear: rather than transferring as a shared set of biases, harmful associations adapt culturally to the prompt language and amplify bias against locally salient protected groups. \textbf{(iii)} Human and LLM harmfulness judgments are broadly aligned (Spearman $ρ=0.62$), with disagreements concentrating on specific attribute classes rather than specific providers. To support further analyses, we release the evaluation code and the dataset, including model generations, attribute annotations, and harmfulness ratings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。