让大模型评测更难,能看清不同模型的真实差距。
Enhancing LLM Evaluations: The Garbling Trick
- 把原有评测逐步升级为更难的任务
- 新评测能暴露原评测看不出来的性能差异
- 适合评估模型推理能力,尤其是新旧模型对比
随着大语言模型(LLMs)能力不断增强,传统评估指标趋于饱和,难以区分模型优劣。本文提出一种通用方法,将现有LLM评测转化为一系列逐步增强难度的任务。这些强化后的评估侧重推理能力,能够揭示原始评估中无法察觉的模型相对性能差异。为验证该方法的有效性,我们构建了一个新的多项选择测试语料库,并扩展为一组评估任务,对多种LLM进行了评估。结果揭示了这些模型间的比较能力,尤其凸显了基础模型与近年'推理型'模型之间的差异。
原文摘要 · Abstract (English)
As large language models (LLMs) become increasingly powerful, traditional evaluation metrics tend to saturate, making it challenging to distinguish between models. We propose a general method to transform existing LLM evaluations into a series of progressively more difficult tasks. These enhanced evaluations emphasize reasoning capabilities and can reveal relative performance differences that are not apparent in the original assessments. To demonstrate the effectiveness of our approach, we create a new multiple-choice test corpus, extend it into a family of evaluations, and assess a collection of LLMs. Our results offer insights into the comparative abilities of these models, particularly highlighting the differences between base LLMs and more recent "reasoning" models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。