arXiv:2503.02776cs.CLcs.AI2025-03综述被引 15

梳理大模型隐性偏见的检测与评估方法,揭示其潜在风险。

Implicit Bias in LLMs: A Survey

  • 基于心理实验框架,提出三类隐性偏见检测方法。
  • 构建单值与对比值两类评估指标体系。
  • 适合关注模型公平性与可解释性的研究者阅读。

由于开发者实施了防护机制,大语言模型(LLMs)在显性偏见测试中表现出色。然而,偏见不仅可能显性存在,也可能以隐性形式出现,如同人类虽努力保持公正,仍存在无意识偏见。这种无意识且自动化的特性使隐性偏见研究尤为困难。本文全面综述了现有关于大模型隐性偏见的研究文献。我们首先介绍心理学中与隐性偏见相关的关键概念、理论与方法,并将其延伸至大模型。借鉴内隐联想测试(IAT)等心理框架,我们将检测方法分为三类:词关联、任务导向文本生成和决策行为。评估指标分为单值型与对比值型两类。数据集分为掩码词句与完整句子两类,涵盖多个领域,体现大模型的广泛应用。尽管针对隐性偏见的缓解研究仍有限,本文总结了现有工作并展望未来挑战。旨在为研究人员提供清晰指引,激发创新思路,推动该方向发展。

原文摘要 · Abstract (English)

Due to the implement of guardrails by developers, Large language models (LLMs) have demonstrated exceptional performance in explicit bias tests. However, bias in LLMs may occur not only explicitly, but also implicitly, much like humans who consciously strive for impartiality yet still harbor implicit bias. The unconscious and automatic nature of implicit bias makes it particularly challenging to study. This paper provides a comprehensive review of the existing literature on implicit bias in LLMs. We begin by introducing key concepts, theories and methods related to implicit bias in psychology, extending them from humans to LLMs. Drawing on the Implicit Association Test (IAT) and other psychological frameworks, we categorize detection methods into three primary approaches: word association, task-oriented text generation and decision-making. We divide our taxonomy of evaluation metrics for implicit bias into two categories: single-value-based metrics and comparison-value-based metrics. We classify datasets into two types: sentences with masked tokens and complete sentences, incorporating datasets from various domains to reflect the broad application of LLMs. Although research on mitigating implicit bias in LLMs is still limited, we summarize existing efforts and offer insights on future challenges. We aim for this work to serve as a clear guide for researchers and inspire innovative ideas to advance exploration in this task.

隐性偏见大模型评估方法公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。