新基准测试挑战大模型推理极限,揭示当前能力仍有巨大差距。
BIG-Bench Extra Hard
- 用更难任务替换原BIG-Bench的题目,提升评估难度
- 顶尖通用模型平均准确率仅9.8%,专用模型也仅44.8%
- 适合研究大模型推理瓶颈与评测方法的学者参考
大语言模型在日常应用中日益普及,亟需具备稳健的通用推理能力和多样化的推理技能。然而,现有推理评测主要聚焦数学与编程能力,忽视了更广泛的推理素养。尽管BIG-Bench提供了涵盖多种技能的综合性评估框架,但随着模型进步,其难度已显不足,顶尖模型在较难版本BBH上接近满分,失去区分度。为此,我们提出新的基准BIG-Bench Extra Hard(BBEH),将原BBH中的每个任务替换为更具挑战性的同类任务,以更严格地评估推理能力。我们在BBEH上测试多个模型,发现最佳通用模型的平均准确率为9.8%,最佳推理专用模型为44.8%,表明大模型在通用推理方面仍存在显著提升空间。相关数据集已开源:https://github.com/google-deepmind/bbeh。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly deployed in everyday applications, demanding robust general reasoning capabilities and diverse reasoning skillset. However, current LLM reasoning benchmarks predominantly focus on mathematical and coding abilities, leaving a gap in evaluating broader reasoning proficiencies. One particular exception is the BIG-Bench dataset, which has served as a crucial benchmark for evaluating the general reasoning capabilities of LLMs, thanks to its diverse set of challenging tasks that allowed for a comprehensive assessment of general reasoning across various skills within a unified framework. However, recent advances in LLMs have led to saturation on BIG-Bench, and its harder version BIG-Bench Hard (BBH). State-of-the-art models achieve near-perfect scores on many tasks in BBH, thus diminishing its utility. To address this limitation, we introduce BIG-Bench Extra Hard (BBEH), a new benchmark designed to push the boundaries of LLM reasoning evaluation. BBEH replaces each task in BBH with a novel task that probes a similar reasoning capability but exhibits significantly increased difficulty. We evaluate various models on BBEH and observe a (harmonic) average accuracy of 9.8\% for the best general-purpose model and 44.8\% for the best reasoning-specialized model, indicating substantial room for improvement and highlighting the ongoing challenge of achieving robust general reasoning in LLMs. We release BBEH publicly at: https://github.com/google-deepmind/bbeh.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。