arXiv:2601.03986cs.CL2026-01被引 14

提出评估大模型评测集质量的系统方法,发现现有评测存在显著差异。

Benchmark^2: Systematic Evaluation of LLM Benchmarks

  • 设计三项互补指标评估评测集质量:排名一致性、区分度和能力偏差。
  • 15个评测集对比显示不同评测结果差异大,部分评测结果不可靠。
  • 按指标筛选评测可大幅减少测试数据量,同时保持评估效果。

大语言模型评测集的快速增多带来了评估评测集自身质量的迫切需求。我们提出Benchmark^2,一个包含三项互补指标的综合框架:(1) 跨评测排名一致性,衡量评测是否与同行评测结果一致;(2) 区分度得分,量化评测区分不同模型的能力;(3) 能力对齐偏差,识别强模型失败而弱模型成功的问题实例。我们在涵盖数学、推理和知识领域的15个评测集上进行了广泛实验,评估了4个模型家族中的11个LLM。分析发现现有评测集质量差异显著,并表明基于我们的指标进行有选择性的评测构建,可在大幅缩减测试集规模的同时实现相当的评估性能。

原文摘要 · Abstract (English)

The rapid proliferation of benchmarks for evaluating large language models (LLMs) has created an urgent need for systematic methods to assess benchmark quality itself. We propose Benchmark^2, a comprehensive framework comprising three complementary metrics: (1) Cross-Benchmark Ranking Consistency, measuring whether a benchmark produces model rankings aligned with peer benchmarks; (2) Discriminability Score, quantifying a benchmark's ability to differentiate between models; and (3) Capability Alignment Deviation, identifying problematic instances where stronger models fail but weaker models succeed within the same model family. We conduct extensive experiments across 15 benchmarks spanning mathematics, reasoning, and knowledge domains, evaluating 11 LLMs across four model families. Our analysis reveals significant quality variations among existing benchmarks and demonstrates that selective benchmark construction based on our metrics can achieve comparable evaluation performance with substantially reduced test sets.

评测框架大模型评估基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。