arXiv:2507.11216cs.CL2025-07被引 1

构建西语与加泰罗尼亚语问答偏见评估基准,揭示大模型在西班牙社会语境中的偏见问题。

EsBBQ and CaBBQ: The Spanish and Catalan Bias Benchmarks for Question Answering

  • 基于原BBQ构建西语和加泰罗尼亚语平行偏见评估数据集
  • 多模型测试显示高准确率常伴随更强社会偏见依赖
  • 首次系统评估西班牙语场景下大模型的社会偏见

现有研究普遍表明,大型语言模型会继承其预训练数据中的社会偏见。由于非英语语言及美国以外社会背景的偏见评估资源严重不足,本文提出西班牙语和加泰罗尼亚语问答偏见基准(EsBBQ 和 CaBBQ)。这两个并行数据集基于原始 BBQ 构建,采用多项选择题形式,在西班牙社会语境下评估10类社会偏见。我们对多种大模型进行了评估,涵盖模型家族、规模和变体。结果表明,模型在模糊情境下难以选出正确答案,且高问答准确率往往与更强的社会偏见依赖相关。

原文摘要 · Abstract (English)

Previous literature has largely shown that Large Language Models (LLMs) perpetuate social biases learnt from their pre-training data. Given the notable lack of resources for social bias evaluation in languages other than English, and for social contexts outside of the United States, this paper introduces the Spanish and the Catalan Bias Benchmarks for Question Answering (EsBBQ and CaBBQ). Based on the original BBQ, these two parallel datasets are designed to assess social bias across 10 categories using a multiple-choice QA setting, now adapted to the Spanish and Catalan languages and to the social context of Spain. We report evaluation results on different LLMs, factoring in model family, size and variant. Our results show that models tend to fail to choose the correct answer in ambiguous scenarios, and that high QA accuracy often correlates with greater reliance on social biases.

偏见评估多语言问答系统LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。