arXiv:2506.17111cs.AIcs.CL2025-06中稿 · ACL被引 4

不同偏见评估方法对同一模型排名差异大,警示基准测试可靠性问题。

Are Bias Evaluation Methods Biased ?

  • 用多种主流方法评估模型偏见,比较排名一致性
  • 不同方法对同一模型给出显著不同的偏见排序
  • 提醒研究者谨慎使用单一评估基准,需多角度验证

构建用于评估大语言模型安全性的基准是可信AI领域的重要工作。这些基准使模型在毒性、偏见、有害行为等方面得以比较。独立基准采用不同数据集和评估方法。本文通过多种方法对代表性模型进行偏见排名,考察排名的一致性。结果表明,尽管广泛使用,不同偏见评估方法仍导致模型排名差异显著。研究最后为社区提供基准使用建议。

原文摘要 · Abstract (English)

The creation of benchmarks to evaluate the safety of Large Language Models is one of the key activities within the trusted AI community. These benchmarks allow models to be compared for different aspects of safety such as toxicity, bias, harmful behavior etc. Independent benchmarks adopt different approaches with distinct data sets and evaluation methods. We investigate how robust such benchmarks are by using different approaches to rank a set of representative models for bias and compare how similar are the overall rankings. We show that different but widely used bias evaluations methods result in disparate model rankings. We conclude with recommendations for the community in the usage of such benchmarks.

大模型安全偏见评估基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。