arXiv:2502.19749cs.CL2025-02EMNLP被引 3

用真实语境评估大模型隐性社会偏见,发现模型表面无偏实则藏偏。

What's Not Said Still Hurts: A Description-Based Evaluation Framework for Measuring Social Bias in LLMs

  • 设计自然语境下的描述型评测集,捕捉隐藏在上下文中的偏见
  • 测试6个主流大模型,发现其在细微场景中仍持续强化偏见
  • 适合关注模型公平性与隐性偏见检测的研究者和开发者

大型语言模型(LLMs)常继承训练数据中的社会偏见。现有基准通过直接词项关联(如人口术语与偏见词汇)评估偏见,但随着模型能力提升,其在显性词项层面的偏见表现降低,导致传统评测低估真实偏见水平。本文提出描述型偏见评测基准(DBB),一个新型数据集,旨在评估语义层面的隐性偏见——即偏见概念嵌入真实世界、微妙且非显性的自然语境中,而非仅依赖表面词项。我们分析了六种最先进的大模型,结果表明:尽管模型在词项层面的偏见响应显著降低,但在复杂、细微的语境中仍持续表现出系统性偏见。相关数据、代码与结果已公开于 https://github.com/JP-25/Description-based-Bias-Benchmark。

原文摘要 · Abstract (English)

Large Language Models (LLMs) often exhibit social biases inherited from their training data. While existing benchmarks evaluate bias by term-based mode through direct term associations between demographic terms and bias terms, LLMs have become increasingly adept at avoiding biased responses, leading to seemingly low levels of bias. However, biases persist in subtler, contextually hidden forms that traditional benchmarks fail to capture. We introduce the Description-based Bias Benchmark (DBB), a novel dataset designed to assess bias at the semantic level that bias concepts are hidden within naturalistic, subtly framed contexts in real-world scenarios rather than superficial terms. We analyze six state-of-the-art LLMs, revealing that while models reduce bias in response at the term level, they continue to reinforce biases in nuanced settings. Data, code, and results are available at https://github.com/JP-25/Description-based-Bias-Benchmark.

社会偏见大模型评估隐性偏见评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。