arXiv:2507.02088cs.CL2025-07ACL被引 4

构建首个中文多任务偏见评估基准,覆盖12类偏见

McBE: A Multi-task Chinese Bias Evaluation Benchmark for Large Language Models

  • 设计涵盖12类偏见的多任务评估框架
  • 包含4077个样本,覆盖82个子类别
  • 适合研究中文LLM偏见与伦理风险的学者

随着大语言模型(LLMs)在各类自然语言处理任务中广泛应用,其内在偏见逐渐暴露。因此,衡量LLMs中的偏见对降低伦理风险至关重要。然而,现有偏见评估数据集多聚焦于英语和北美文化,其偏见类别难以适用于其他文化。针对中文语境的数据集尤为稀缺。更重要的是,这些数据集通常仅支持单一评估任务,无法从多维度评估LLMs的偏见。为此,我们提出了一个面向中文的多任务偏见评估基准(McBE),包含4,077个偏见评估实例,覆盖12个单类偏见、82个子类别,并引入5个评估任务,实现广泛的类别覆盖、内容多样性和评估全面性。此外,我们评估了多个系列及不同参数规模的主流LLMs。总体而言,所有测试模型均表现出不同程度的偏见。我们对结果进行了深入分析,为理解LLMs中的偏见提供了新视角。

原文摘要 · Abstract (English)

As large language models (LLMs) are increasingly applied to various NLP tasks, their inherent biases are gradually disclosed. Therefore, measuring biases in LLMs is crucial to mitigate its ethical risks. However, most existing bias evaluation datasets focus on English and North American culture, and their bias categories are not fully applicable to other cultures. The datasets grounded in the Chinese language and culture are scarce. More importantly, these datasets usually only support single evaluation tasks and cannot evaluate the bias from multiple aspects in LLMs. To address these issues, we present a Multi-task Chinese Bias Evaluation Benchmark (McBE) that includes 4,077 bias evaluation instances, covering 12 single bias categories, 82 subcategories and introducing 5 evaluation tasks, providing extensive category coverage, content diversity, and measuring comprehensiveness. Additionally, we evaluate several popular LLMs from different series and with parameter sizes. In general, all these LLMs demonstrated varying degrees of bias. We conduct an in-depth analysis of results, offering novel insights into bias in LLMs.

偏见评估中文LLM多任务基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。