arXiv:2512.06732cs.CLcs.AI2025-12被引 2

新基准揭示大模型隐性偏见,比显性测试更难发现。

"The Dentist is an involved parent, the bartender is not": Revealing Implicit Biases in QA with Implicit BBQ

  • 用文化线索和姓名暗示身份,模拟真实世界隐性偏见
  • GPT-4o在性取向类任务中准确率下降最高达7%
  • 适合关注模型公平性与社会影响的研究者

现有大语言模型偏见评估基准多依赖显式提示,如直接标明宗教、种族、性别等身份。然而现实互动中的偏见常通过姓名、文化线索或特质间接体现,此类隐性偏见未被充分覆盖,造成公平性评估的重大盲区。本文提出ImplicitBBQ,扩展了偏见问答基准BBQ,涵盖6个类别中的隐性提示身份属性。对GPT-4o的评估显示,相比显性提示,其在‘性取向’子类别中准确率下降高达7%,且多数类别均出现一致下降。这表明当前大模型存在显式基准无法检测的隐性偏见。ImplicitBBQ为自然语言处理中的精细公平性评估提供了关键工具。

原文摘要 · Abstract (English)

Existing benchmarks evaluating biases in large language models (LLMs) primarily rely on explicit cues, declaring protected attributes like religion, race, gender by name. However, real-world interactions often contain implicit biases, inferred subtly through names, cultural cues, or traits. This critical oversight creates a significant blind spot in fairness evaluation. We introduce ImplicitBBQ, a benchmark extending the Bias Benchmark for QA (BBQ) with implicitly cued protected attributes across 6 categories. Our evaluation of GPT-4o on ImplicitBBQ illustrates troubling performance disparity from explicit BBQ prompts, with accuracy declining up to 7% in the "sexual orientation" subcategory and consistent decline located across most other categories. This indicates that current LLMs contain implicit biases undetected by explicit benchmarks. ImplicitBBQ offers a crucial tool for nuanced fairness evaluation in NLP.

模型偏见公平性评估隐性偏见

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。