arXiv:2602.14466cs.CL2026-02中稿 · LREC 2026

构建菲律宾语偏见测试集,评估大模型在本地语境下的性别与性向偏见。

Robust Bias Evaluation with FilBBQ: A Filipino Bias Benchmark for Question-Answering Language Models

  • 通过四阶段流程构建覆盖1万+提示的菲律宾语偏见测试集
  • 发现模型在情感、家庭角色等维度存在显著性别与性向偏见
  • 采用多随机种子平均提升评估可靠性,适合本地化语言模型评测

随着自然语言生成应用的普及,问答式偏见基准(BBQ)已成为评估生成模型刻板印象的重要工具。本文拓展了BBQ的语言范围,通过模板分类、文化敏感翻译、新模板构建和提示生成四个阶段,构建了针对菲律宾语境的FilBBQ偏见测试集。该测试集包含超过10,000个提示,用于评估模型在菲律宾社会背景下是否表现出性别歧视与同性恋偏见。我们采用改进的稳健评估协议,在多个随机种子下获取模型响应并取平均值,以应对响应不稳定性问题。结果表明,不同种子间偏见分数存在显著差异,且模型在情感表达、家庭角色、刻板化酷儿兴趣及一夫多妻制等方面仍存在明显偏见。FilBBQ已开源:https://github.com/gamboalance/filbbq。

原文摘要 · Abstract (English)

With natural language generation becoming a popular use case for language models, the Bias Benchmark for Question-Answering (BBQ) has grown to be an important benchmark format for evaluating stereotypical associations exhibited by generative models. We expand the linguistic scope of BBQ and construct FilBBQ through a four-phase development process consisting of template categorization, culturally aware translation, new template construction, and prompt generation. These processes resulted in a bias test composed of more than 10,000 prompts which assess whether models demonstrate sexist and homophobic prejudices relevant to the Philippine context. We then apply FilBBQ on models trained in Filipino but do so with a robust evaluation protocol that improves upon the reliability and accuracy of previous BBQ implementations. Specifically, we account for models' response instability by obtaining prompt responses across multiple seeds and averaging the bias scores calculated from these distinctly seeded runs. Our results confirm both the variability of bias scores across different seeds and the presence of sexist and homophobic biases relating to emotion, domesticity, stereotyped queer interests, and polygamy. FilBBQ is available via https://github.com/gamboalance/filbbq.

偏见评估语言模型菲律宾语多种子评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。