动态调整难度的多学科推理评估基准,适配大模型进化能力。
MorphoBench: A Benchmark with Difficulty Adaptive to Model Reasoning
- 基于模型推理过程生成关键陈述,动态调节题目难度。
- 涵盖1300+道奥赛级复杂推理题,支持O3/GPT-5等模型迭代测试。
- 适用于评估与提升大模型推理能力,尤其适合研发人员使用。
随着强大大型推理模型的发展,有效评估其推理能力变得日益重要。然而,现有用于评估大型模型推理能力的基准测试往往范围有限,且难以根据模型推理能力的演进动态调整难度。为此,我们提出MorphoBench,一个包含多学科问题的基准测试,能够根据先进模型的推理能力自适应地调整和更新题目难度。具体而言,我们从奥赛级竞赛等现有基准和来源中筛选并收集复杂推理题;同时,通过利用模型推理过程中生成的关键陈述,动态调整题目的分析挑战性;此外,还引入模拟软件生成的问题,实现低资源消耗的难度动态调整。我们已收集超过1,300个测试题,并基于O3和GPT-5等模型的推理能力对MorphoBench进行了迭代难度调整。该基准显著提升了模型推理评估的全面性和有效性,为提升大型模型的推理能力与科学稳健性提供了可靠指导。代码已开源:https://github.com/OpenDCAI/MorphoBench。
原文摘要 · Abstract (English)
With the advancement of powerful large-scale reasoning models, effectively evaluating the reasoning capabilities of these models has become increasingly important. However, existing benchmarks designed to assess the reasoning abilities of large models tend to be limited in scope and lack the flexibility to adapt their difficulty according to the evolving reasoning capacities of the models. To address this, we propose MorphoBench, a benchmark that incorporates multidisciplinary questions to evaluate the reasoning capabilities of large models and can adjust and update question difficulty based on the reasoning abilities of advanced models. Specifically, we curate the benchmark by selecting and collecting complex reasoning questions from existing benchmarks and sources such as Olympiad-level competitions. Additionally, MorphoBench adaptively modifies the analytical challenge of questions by leveraging key statements generated during the model's reasoning process. Furthermore, it includes questions generated using simulation software, enabling dynamic adjustment of benchmark difficulty with minimal resource consumption. We have gathered over 1,300 test questions and iteratively adjusted the difficulty of MorphoBench based on the reasoning capabilities of models such as o3 and GPT-5. MorphoBench enhances the comprehensiveness and validity of model reasoning evaluation, providing reliable guidance for improving both the reasoning abilities and scientific robustness of large models. The code has been released in https://github.com/OpenDCAI/MorphoBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。