发现大模型隐含的刻板印象轴,揭示其与人类认知的差异。
STEREODISCO: Discovering Stereotypicality in LLMs

- 用语义差异法构建2000个候选刻板轴,探测模型内部表征
- 发现模型对社会群体的刻板评价与人类不一致,且存在新轴
- 识别出谦逊/傲慢等未被研究的刻板维度,获人工验证
大语言模型会编码、传递并强化刻板印象。以往计算研究仅聚焦社会心理学中的少数语义轴,且基于模型生成的词向量,未揭示其他潜在的刻板语义轴及其在模型内部的表征方式。本文提出 STEREODISCO 框架,将语义差异法(Osgood et al., 1957) adapted 用于系统研究模型内部表征中的刻板印象。该框架从 WordNet 反义词集构造约 2,000 个候选语义轴,通过探针技术在模型激活空间中恢复每个轴,并基于概念投影的统计检验识别刻板轴。以 LLAMA-3-8B-INSTRUCT 与 MISTRAL-7B-INSTRUCT 为例进行案例研究,发现两模型对社会群体的刻板评分相关性高于与人类的一致性,表明模型所编码的刻板内容与社会心理学文献记录不同。此外,还发现了此前未被研究的刻板轴——如谦逊 vs 傲慢、狭隘思维 vs 广阔思维、怯懦 vs 勇敢——经人工标注独立验证为真实存在的刻板维度。
原文摘要 · Abstract (English)
LLMs encode, convey, and perpetuate stereotypes. Prior computational research focuses on a small set of semantic axes investigated in social psychology, and operates on word embeddings produced by language models, leaving open which other semantic axes carry stereotypical associations in LLMs and how LLMs internally represent such axes. We introduce STEREODISCO, a framework that adapts the semantic differential method (Osgood et al., 1957) to the systematic study of stereotypes in LLM internal representations. STEREODISCO constructs approx. 2,000 candidate semantic axes from WordNet antonym synsets, recovers each as a geometric axis in the LLM's activation space via probing, and identifies stereotypical axes via a statistical test over concept projections. As a case study, we apply STEREODISCO to social group stereotypes with LLAMA-3-8B-INSTRUCT and MISTRAL-7B-INSTRUCT. We find that the two LLMs agree with each other on social group ratings more than with humans, suggesting that LLM-encoded stereotype content diverges from that documented in social psychology. We also discover stereotypical axes not investigated in prior work -- including humble vs. proud, narrow-minded vs. broad-minded, and cowardly vs. brave, which human annotators independently confirm.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。