arXiv:2510.12857cs.CYcs.AI2025-10

自动生成真实对话式问题,更精准检测大模型偏见。

Adaptive Generation of Bias-Eliciting Questions for LLMs

  • 用反事实迭代法生成开放问答,模拟真实用户交互。
  • 发现所有测试模型在特定场景下仍存在持续偏见。
  • 适合做模型公平性评估的研究者和开发者使用。

大语言模型已广泛应用于面向用户的应用中,覆盖数亿用户。尽管应用广泛,但对输出依赖的增加引发了对其固有偏见的担忧,这些偏见可能损害或刻板化特定群体。现有偏见评测基准多依赖简单模板提示或限制性选择题,无法反映真实用户交互的复杂性。本文提出一种反事实框架,可自动生成真实、开放的问题用于大模型偏见评估。通过迭代问题变异,该方法系统探索模型最易出现偏见的区域。除检测有害偏见外,还捕捉了日益重要的响应维度,如不对称拒绝和显式偏见承认。基于此,我们构建了CAB——一个多样化且经人工验证的基准,用于当前前沿大模型的现实与细致偏见评估。使用CAB的评估表明,所有被检模型在某些场景下仍存在持续偏见,凸显公平性研究的必要性。

原文摘要 · Abstract (English)

Large language models (LLMs) are now widely deployed in user-facing applications, reaching hundreds of millions of users worldwide. Despite their widespread adoption, growing reliance on their outputs raises significant concerns, particularly as users may be exposed to model-inherent biases that disadvantage or stereotype certain groups. However, existing bias benchmarks commonly rely on simple templated prompts or restrictive multiple-choice questions that fail to capture the complexity of real-world user interactions. In this work, we address this gap by introducing a counterfactual framework that automatically generates realistic, open-ended questions for LLM bias evaluation. Through iterative question mutation, our approach systematically explores areas where models are most likely to exhibit biased behavior. Beyond just detecting harmful biases, we also capture increasingly relevant response dimensions, such as asymmetric refusals and explicit bias acknowledgment. Building on this, we construct CAB, a diverse and human-verified benchmark for realistic and nuanced bias evaluations on current frontier LLMs. Our evaluation using CAB highlights the continued need for fairness research by showing that all examined models exhibit persistent biases across certain scenarios.

大模型偏见评测基准自动化评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。