为大模型推理能力评估提供6个带难度分的标准化数据集
Easy2Hard-Bench: Standardized Difficulty Labels for Profiling LLM Performance and Generalization
- 用真实人类和大模型答题数据,结合IRT等模型标定题目难度
- 覆盖数学、编程等多领域,难题占比高于以往数据集
- 适合研究大模型泛化能力或做难度敏感性分析的研究者
当前缺乏在广泛复杂度范围内对问题进行细粒度难度标注的数据集。为此,我们提出Easy2Hard-Bench,一个包含6个跨领域(如数学、编程、国际象棋谜题、推理题)基准数据集的统一格式集合。每个问题均配有数值化难度评分。通过收集真实人类或主流大模型在权威榜单上的答题表现数据,利用项目反应理论(IRT)和Glicko-2等成熟难度评估模型,对问题进行统一打分。相较于已有数据集,Easy2Hard-Bench包含更高比例的挑战性问题。基于六种先进大模型的实验表明,该数据集能全面刻画模型在不同难度下的性能与泛化能力,推动未来大模型泛化研究。数据集已开源:https://huggingface.co/datasets/furonghuang-lab/Easy2Hard-Bench。
原文摘要 · Abstract (English)
While generalization over tasks from easy to hard is crucial to profile language models (LLMs), the datasets with fine-grained difficulty annotations for each problem across a broad range of complexity are still blank. Aiming to address this limitation, we present Easy2Hard-Bench, a consistently formatted collection of 6 benchmark datasets spanning various domains, such as mathematics and programming problems, chess puzzles, and reasoning questions. Each problem within these datasets is annotated with numerical difficulty scores. To systematically estimate problem difficulties, we collect abundant performance data on attempts to each problem by humans in the real world or LLMs on the prominent leaderboard. Leveraging the rich performance data, we apply well-established difficulty ranking systems, such as Item Response Theory (IRT) and Glicko-2 models, to uniformly assign numerical difficulty scores to problems. Moreover, datasets in Easy2Hard-Bench distinguish themselves from previous collections by a higher proportion of challenging problems. Through extensive experiments with six state-of-the-art LLMs, we provide a comprehensive analysis of their performance and generalization capabilities across varying levels of difficulty, with the aim of inspiring future research in LLM generalization. The datasets are available at https://huggingface.co/datasets/furonghuang-lab/Easy2Hard-Bench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。