arXiv:2510.08730cs.CLcs.LG2025-10被引 3

微基准测试常不可靠,小样本评估易失真。

How Reliable is Language Model Micro-Benchmarking?

  • 提出元评估指标,衡量微基准对模型排序的可靠性
  • 250个样本才可稳定区分性能相差3.5分的模型
  • 25个样本无法保证8B模型对比的可信性

微基准测试为语言模型开发提供高效评估方案:仅用少量现有基准数据即可完成评测。但这些微基准能否像完整基准一样可靠地排名模型?是否优于随机采样?我们发现,在多数情况下答案是否定的。本文提出一种元评估方法,分析微基准在不同模型性能差距下正确排序的能力。研究显示,要稳定区分在MMLU-Pro上相差3.5分或BIG-bench Hard上相差4分的模型对,需至少250个样本;而25个样本时,超过一半的8B指令微调模型比较结果不可靠。当样本量达250时,随机采样已与现有微基准方法表现相当。本研究为评估效率与可靠性之间的权衡提供了实用指导。

原文摘要 · Abstract (English)

Micro-benchmarking offers a solution to the often prohibitive time and cost of language model development: evaluate on a very small subset of existing benchmarks. Can these micro-benchmarks, however, rank models as consistently as the full benchmarks they replace? And can they rank models more consistently than selecting a random subset of data points? In many scenarios, we find that the answer is no. We introduce a meta-evaluation measure for micro-benchmarking which investigates how well a micro-benchmark can rank two models as a function of their performance difference on the full benchmark. This approach can determine which model pairs can be ranked correctly by a micro-benchmark, allowing for a finer-grained analysis of the trade-off between micro-benchmark size and reliability. Prior work has suggested selecting as few as 10 examples; we find that no micro-benchmarking method can consistently rank model pairs 3.5 points of accuracy apart on MMLU-Pro or 4 points apart on BIG-bench Hard. In order to consistently rank model pairs with relatively similar performances, we show that often as many as 250 examples must be selected, at which point random sampling is competitive with existing micro-benchmarking methods. When comparing only 8B instruction-tuned models on MMLU-Pro micro-benchmarks with 25 examples, we find that more than half of pairwise comparisons are not likely to be preserved. Our work provides actionable guidance for both micro-benchmark users and developers in navigating the trade-off between evaluation efficiency and reliability.

模型评估微基准可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。