arXiv:2510.04641cs.CLcs.CY2025-10被引 3

评测大模型检测社会偏见能力,发现小模型微调后效果更好但仍有短板。

Evaluating LLMs for Demographic-Targeted Social Bias Detection: A Comprehensive Benchmark Study

  • 构建多标签偏见检测任务,覆盖多种身份和内容类型。
  • 微调的小模型在多数数据集上表现优于大模型,适合规模化应用。
  • 跨身份联合偏见仍难识别,需更优检测框架。

用于训练通用AI模型的大规模网络爬取文本常包含有害的针对特定群体的社会偏见,引发监管层面的数据审计需求,并推动可扩展的偏见检测方法发展。尽管已有研究探讨过文本数据集中的偏见及检测方法,但这些研究范围有限:通常仅关注单一内容类型(如仇恨言论),覆盖少数人口统计维度,忽略多重群体同时受影响的偏见,且分析技术种类较少。因此,从业者缺乏对近期大语言模型(LLMs)在自动化偏见检测中优势与局限的全面理解。本研究针对英文文本开展综合性基准评估,系统检验LLMs在检测人口统计目标型社会偏见方面的能力。为契合监管要求,我们将偏见检测定义为使用以人口统计为中心的分类体系进行多标签身份识别的任务。我们跨模型规模与技术(包括提示、上下文学习、微调)进行评估。基于十二个涵盖多样内容类型与人口维度的数据集,研究显示微调后的较小模型在可扩展性检测中具有潜力。然而分析也揭示了在不同人口轴线上持续存在的差距,以及对多重群体靶向偏见的识别不足,凸显出开发更有效、可扩展检测框架的必要性。

原文摘要 · Abstract (English)

Large-scale web-scraped text corpora used to train general-purpose AI models often contain harmful demographic-targeted social biases, creating a regulatory need for data auditing and developing scalable bias-detection methods. Although prior work has investigated biases in text datasets and related detection methods, these studies remain narrow in scope. They typically focus on a single content type (e.g., hate speech), cover limited demographic axes, overlook biases affecting multiple demographics simultaneously, and analyze limited techniques. Consequently, practitioners lack a holistic understanding of the strengths and limitations of recent large language models (LLMs) for automated bias detection. In this study, we conduct a comprehensive benchmark study on English texts to assess the ability of LLMs in detecting demographic-targeted social biases. To align with regulatory requirements, we frame bias detection as a multi-label task of detecting targeted identities using a demographic-focused taxonomy. We then systematically evaluate models across scales and techniques, including prompting, in-context learning, and fine-tuning. Using twelve datasets spanning diverse content types and demographics, our study demonstrates the promise of fine-tuned smaller models for scalable detection. However, our analyses also expose persistent gaps across demographic axes and multi-demographic targeted biases, underscoring the need for more effective and scalable detection frameworks.

偏见检测大模型评估社会影响

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。