模型能力饱和,现有评测已难真实反映推理水平。
The Ouroboros of Benchmarking: Reasoning Evaluation in an Era of Saturation
- 分析三大厂商模型在多年间的评测表现变化
- 发现多数基准测试结果趋于饱和,难以区分模型优劣
- 呼吁重新思考评测设计,适合评估与模型开发研究者
大型语言模型(LLMs)和大型推理模型(LRMs)的快速发展伴随了评测基准数量的激增。然而,由于模型能力随规模提升及训练方法创新,加上许多数据集可能已被用于预训练或后训练,导致性能逐渐饱和,催生对更难、更新基准的持续需求。本文聚焦OpenAI、Anthropic、Google三大模型家族,分析其推理能力在不同基准上的演进趋势。通过考察多年间各类推理任务的表现变化,揭示当前评测体系的局限性与挑战。本研究旨在为未来推理评估与模型研发提供首个综合性参考框架。
原文摘要 · Abstract (English)
The rapid rise of Large Language Models (LLMs) and Large Reasoning Models (LRMs) has been accompanied by an equally rapid increase of benchmarks used to assess them. However, due to both improved model competence resulting from scaling and novel training advances as well as likely many of these datasets being included in pre or post training data, results become saturated, driving a continuous need for new and more challenging replacements. In this paper, we discuss whether surpassing a benchmark truly demonstrates reasoning ability or are we simply tracking numbers divorced from the capabilities we claim to measure? We present an investigation focused on three model families, OpenAI, Anthropic, and Google, and how their reasoning capabilities across different benchmarks evolve over the years. We also analyze performance trends over the years across different reasoning tasks and discuss the current situation of benchmarking and remaining challenges. By offering a comprehensive overview of benchmarks and reasoning tasks, our work aims to serve as a first reference to ground future research in reasoning evaluation and model development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。