为大模型评测体系建立健康度评分,帮研究者选可信基准。
Benchmark Health Index: A Systematic Framework for Benchmarking the Benchmarks of LLMs
- 从能力区分、抗饱和性、影响力三维度量化评测集健康度。
- 基于91个模型的106个基准分析,发现多数基准已出现分数虚高。
- 适合评估模型性能的研究者和评测系统设计者使用。
大语言模型快速发展,但用于衡量进展的评测标准正变得不可靠。分数虚高与选择性报告削弱了主流评测的公信力,令学界难以判断哪些结果仍可信。我们提出基准健康指数(BHI),一个纯数据驱动的框架,从三个正交互补维度审计评测集:(1) 能力区分度,衡量评测能否清晰区分模型性能超越噪声水平;(2) 抗饱和性,估算达到天花板前的剩余提升空间,以评估评测的长期有效性;(3) 影响力,通过学术与产业生态中的采纳广度及实践引导力来量化。通过对2025年91个代表性模型的技术报告中提取的106个经验证基准进行系统分析,我们首次在宏观层面量化了评测体系的健康状况,为基准选择提供理论依据,并支持下一代评测协议的动态生命周期管理。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are advancing rapidly, yet the benchmarks used to measure this progress are becoming increasingly unreliable. Score inflation and selective reporting have eroded the authority of standard benchmarks, leaving the community uncertain about which evaluation results remain trustworthy. We introduce the Benchmark Health Index (BHI), a pure data-driven framework for auditing evaluation sets along three orthogonal and complementary axes: (1) Capability Discrimination, measuring how sharply a benchmark separates model performance beyond noise; (2) Anti-Saturation, estimating remaining headroom before ceiling effects erode resolution and thus the benchmark's expected longevity; and (3) Impact, quantifying influence across academic and industrial ecosystems via adoption breadth and practice-shaping power. By distilling 106 validated benchmarks from the technical reports of 91 representative models in 2025, we systematically characterize the evaluation landscape. BHI is the first framework to quantify benchmark health at a macro level, providing a principled basis for benchmark selection and enabling dynamic lifecycle management for next-generation evaluation protocols.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。