用连续解题评估大模型学习能力,发现强模型未必会进步。
EvaLearn: Quantifying the Learning Capability and Efficiency of LLMs via Sequential Problem Solving
- 设计连续解题任务,让模型从过往经验中学习。
- 9个前沿模型表现各异,部分初始一般但进步快。
- 适合研究模型动态学习能力的学者使用。
我们提出EvaLearn,首个专门评估大语言模型在复杂任务中学习能力与效率的基准。该基准包含648道难题,分属六类任务,组成182个序列,要求模型按顺序求解,从而利用先前经验。不同于多数并行评估,EvaLearn通过五项自动化指标衡量学习进展。实验评测了九个前沿模型,发现如Claude-3.7-sonnet虽初始表现中等,但学习能力强;而部分模型无法受益于经验,甚至出现负迁移。进一步研究显示,实例级评分与教师模型反馈可促进学习。值得注意的是,静态能力更强的模型在学习能力上并无普遍优势,表明EvaLearn揭示了模型性能的新维度。所有数据集、评估框架及结果已开源。
原文摘要 · Abstract (English)
We introduce EvaLearn, a pioneering benchmark designed to evaluate large language models (LLMs) on their learning capability and efficiency in challenging tasks, a critical, yet underexplored aspect of model potential. EvaLearn contains 648 challenging problems across six task types, grouped into 182 sequences, each sequence dedicated to one task type. Diverging from most existing benchmarks that evaluate models in parallel, EvaLearn requires models to solve problems sequentially, allowing them to leverage the experience gained from previous solutions. EvaLearn provides five comprehensive automated metrics to evaluate models and quantify their learning capability and efficiency. We extensively benchmark nine frontier models and observe varied performance profiles: some models, such as Claude-3.7-sonnet, start with moderate initial performance but exhibit strong learning ability, while some models struggle to benefit from experience and may even show negative transfer. Moreover, we investigate model performance under two learning settings and find that instance-level rubrics and teacher-model feedback further facilitate model learning. Importantly, we observe that current LLMs with stronger static abilities do not show a clear advantage in learning capability across all tasks, highlighting that EvaLearn evaluates a new dimension of model performance. We hope EvaLearn provides a novel evaluation perspective for assessing LLM potential and understanding the gap between models and human capabilities, promoting the development of deeper and more dynamic evaluation approaches. All datasets, the automatic evaluation framework, and the results studied in this paper are available at the GitHub repository.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。