arXiv:2503.06987cs.CLcs.AI2025-03ACL被引 14

构建生成式社会偏见评测基准,发现生成与问答评估结果不一致。

Social Bias Benchmark for Generation: A Comparison of Generation and QA-Based Evaluations

  • 用故事续写任务评测大模型生成中的社会偏见
  • 在英韩双语下测试10个模型,发现中性与偏见生成概率差异
  • 揭示生成评估与问答评估结果不一致,适合偏见研究者使用

衡量大语言模型(LLMs)中的社会偏见至关重要,但现有评估方法难以有效评估长文本生成中的偏见。我们提出生成式偏见评测基准(BBG),基于问答式偏见评测基准(BBQ)的改编,通过让大模型续写故事提示来评估其在长文本生成中的社会偏见。我们在英语和韩语环境下构建该基准,评估了十个大模型在中性和偏见生成上的概率。同时,我们将生成式评估结果与多项选择式BBQ评估进行对比,发现两种方法产生不一致的结果。

原文摘要 · Abstract (English)

Measuring social bias in large language models (LLMs) is crucial, but existing bias evaluation methods struggle to assess bias in long-form generation. We propose a Bias Benchmark for Generation (BBG), an adaptation of the Bias Benchmark for QA (BBQ), designed to evaluate social bias in long-form generation by having LLMs generate continuations of story prompts. Building our benchmark in English and Korean, we measure the probability of neutral and biased generations across ten LLMs. We also compare our long-form story generation evaluation results with multiple-choice BBQ evaluation, showing that the two approaches produce inconsistent results.

偏见评测生成评估LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。