arXiv:2501.02295cs.CL2025-01ACL被引 38

揭示大模型隐性偏见远强于显性偏见,且两者不一致。

Explicit vs. Implicit: Investigating Social Bias in Large Language Models through Self-Reflection

  • 用自我反思框架分两阶段测隐性和显性偏见
  • 隐性偏见强,显性偏见弱,且随模型变大更严重
  • 当前对齐方法能压显性偏见,难控隐性偏见

大型语言模型(LLMs)在生成内容中表现出各种偏见和刻板印象。尽管已有大量研究关注模型偏见,但多数聚焦于显性偏见,对隐性偏见及其与显性偏见的关系关注不足。本文基于社会心理学理论,提出一种系统性框架,用于探究并比较LLMs中的显性和隐性偏见。我们设计了一种新颖的自省式评估框架,分两阶段进行:首先通过模拟心理测评方法测量隐性偏见,再通过提示模型分析自身生成内容来评估显性偏见。在多个社会维度上对先进LLMs进行广泛实验发现,模型在显性与隐性偏见之间存在显著不一致性:显性偏见表现为轻微刻板印象,而隐性偏见则呈现强烈刻板印象。进一步分析表明,训练数据规模、模型大小和对齐技术是导致这种不一致的关键因素。实验结果显示,随着训练数据量和模型规模增加,显性偏见下降,而隐性偏见反而上升;当代对齐方法虽能有效抑制显性偏见,但在缓解隐性偏见方面作用有限。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have been shown to exhibit various biases and stereotypes in their generated content. While extensive research has investigated biases in LLMs, prior work has predominantly focused on explicit bias, with minimal attention to implicit bias and the relation between these two forms of bias. This paper presents a systematic framework grounded in social psychology theories to investigate and compare explicit and implicit biases in LLMs. We propose a novel self-reflection-based evaluation framework that operates in two phases: first measuring implicit bias through simulated psychological assessment methods, then evaluating explicit bias by prompting LLMs to analyze their own generated content. Through extensive experiments on advanced LLMs across multiple social dimensions, we demonstrate that LLMs exhibit a substantial inconsistency between explicit and implicit biases: while explicit bias manifests as mild stereotypes, implicit bias exhibits strong stereotypes. We further investigate the underlying factors contributing to this explicit-implicit bias inconsistency, examining the effects of training data scale, model size, and alignment techniques. Experimental results indicate that while explicit bias declines with increased training data and model size, implicit bias exhibits a contrasting upward trend. Moreover, contemporary alignment methods effectively suppress explicit bias but show limited efficacy in mitigating implicit bias.

大模型偏见隐性偏见自省评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。