提出新模型提升大模型评估的可靠性,解决基准测试结果与真实能力脱节问题。
Quantifying construct validity in large language model evaluations
- 融合尺度定律与潜在因子模型优势,分离模型规模与真实能力
- 在多个基准上预测表现更准,且解释性更强
- 适合关注模型真实能力评估的研究者和评测工程师
大语言模型社区常将基准测试结果等同于模型通用能力,但测试集污染和标注误差等问题会扭曲性能评估。如何判断一个基准是否可靠反映目标能力?这涉及构建效度问题。现有方法中,潜在因子模型忽略尺度规律,提取的能力常等价于模型大小;尺度定律忽略测量误差,导致能力不可解释且过拟合于已有基准。本文提出结构化能力模型,首次从大量大模型基准结果中提取可解释、泛化性强的能力。在OpenLLM Leaderboard数据上对比实验表明,该模型在简约拟合指标上优于潜在因子模型,在分布外基准预测上优于尺度定律。其优势源于同时考虑模型规模对能力的影响,以及能力通过测量误差影响观测结果的双重机制,显著提升对大模型评估构建效度的量化能力。
原文摘要 · Abstract (English)
The LLM community often reports benchmark results as if they are synonymous with general model capabilities. However, benchmarks can have problems that distort performance, like test set contamination and annotator error. How can we know that a benchmark is a reliable indicator of some capability that we want to measure? This question concerns the construct validity of LLM benchmarks, and it requires separating benchmark results from capabilities when we model and predict LLM performance. Both social scientists and computer scientists propose formal models - latent factor models and scaling laws - for identifying the capabilities underlying benchmark scores. However, neither technique is satisfactory for construct validity. Latent factor models ignore scaling laws, and as a result, the capabilities they extract often proxy model size. Scaling laws ignore measurement error, and as a result, the capabilities they extract are both uninterpretable and overfit to the observed benchmarks. This thesis presents the structured capabilities model, the first model to extract interpretable and generalisable capabilities from a large collection of LLM benchmark results. I fit this model and its two alternatives on a large sample of results from the OpenLLM Leaderboard. Structured capabilities outperform latent factor models on parsimonious fit indices, and exhibit better out-of-distribution benchmark prediction than scaling laws. These improvements are possible because neither existing approach separates model scale from capabilities in the appropriate way. Model scale should inform capabilities, as in scaling laws, and these capabilities should inform observed results up to measurement error, as in latent factor models. In combining these two insights, structured capabilities demonstrate better explanatory and predictive power for quantifying construct validity in LLM evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。