通过训练后再测试,让大模型排名更一致可靠。
Train-before-Test Harmonizes Language Model Rankings
- 先统一微调再评估,揭示模型真实潜力。
- 跨基准排名高度一致,打破传统评估矛盾。
- 仅一个潜在因素决定模型能力,适合选型参考。
现有语言模型基准测试常得出相互矛盾的模型排名,即使针对相似能力。本文提出‘训练后再测试’方法:在评估前对所有模型进行相同基准的微调。实验覆盖24个基准和61个模型。结果表明,该方法下模型潜力排名在各基准间高度一致,具有显著外部有效性;而传统直接评估则缺乏一致性。此外,该方法恢复了困惑度与下游任务性能的关联,预训练困惑度能预测微调后表现,说明排名反映的是内在潜力而非微调偏差。最后,模型评分矩阵可简化为单一主因子,揭示模型潜力由一个核心因素主导。建议将训练后再测试作为大模型评测的标准流程。
原文摘要 · Abstract (English)
Existing language model benchmarks provide contradictory model rankings, even for benchmarks that aim to capture similar skills. This dilemma of conflicting rankings hampers model selection, clouds model comparisons, and adds confusion to a growing ecosystem of competing models. In this paper, we take a different perspective on model comparison: instead of relying on out-of-the-box performance via direct evaluation, we compare model potential by providing each model with identical benchmark-specific fine-tuning before evaluation. We call this approach train-before-test. Our primary contribution is a comprehensive empirical evaluation of model potential across 24 benchmarks and 61 models. First, we demonstrate that model potential rankings obtained through train-before-test exhibit remarkable consistency across all benchmarks. Whereas traditional rankings demonstrate little external validity under direct evaluation, they enjoy a significant degree of external validity when applying train-before-test: model potential rankings transfer gracefully from one benchmark to another. Second, train-before-test restores the connection between perplexity and downstream task performance, lost under direct evaluation. Remarkably, even pre-finetuning perplexity of a base model predicts post-finetuning downstream performance, suggesting that ranking consistency reflects inherent model potential rather than fine-tuning artifacts. Finally, train-before-test reduces the model-score matrix to essentially rank one, indicating that model potential is dominated by one latent factor, uncovered by train-before-test. Our work supports the recommendation to make train-before-test a default component of LLM benchmarking.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。