arXiv:2506.16395cs.CL2025-06被引 19

评测大模型编程竞赛能力,发现顶尖模型仍难应对高难度题目

OJBench: A Competition Level Code Benchmark For Large Language Models

  • 构建232道信息学竞赛题组成的评测基准OJBench
  • 37个模型测试显示顶尖模型在难题上表现依然有限
  • 适合评估大模型代码推理能力,尤其关注竞赛级挑战

大型语言模型(LLMs)在数学与代码推理方面取得显著进展。然而,现有代码评测基准难以全面评估其在竞争级别上的能力。为此,我们提出OJBench,一个全新的、具有挑战性的基准,用于评估LLMs在竞赛级代码推理方面的能力。OJBench包含来自NOI和ICPC的232道编程竞赛题,对模型推理能力提出了更严格的要求。我们在37个模型上进行了全面评估,涵盖闭源与开源、专注推理与非推理类模型。结果显示,即使最先进的推理型模型如o4-mini和Gemini-2.5-pro-exp,在高难度竞赛题上仍表现不佳,凸显了模型在竞赛级代码推理中面临的巨大挑战。

原文摘要 · Abstract (English)

Recent advancements in large language models (LLMs) have demonstrated significant progress in math and code reasoning capabilities. However, existing code benchmark are limited in their ability to evaluate the full spectrum of these capabilities, particularly at the competitive level. To bridge this gap, we introduce OJBench, a novel and challenging benchmark designed to assess the competitive-level code reasoning abilities of LLMs. OJBench comprises 232 programming competition problems from NOI and ICPC, providing a more rigorous test of models' reasoning skills. We conducted a comprehensive evaluation using OJBench on 37 models, including both closed-source and open-source models, reasoning-oriented and non-reasoning-oriented models. Our results indicate that even state-of-the-art reasoning-oriented models, such as o4-mini and Gemini-2.5-pro-exp, struggle with highly challenging competition-level problems. This highlights the significant challenges that models face in competitive-level code reasoning.

代码生成大模型评测竞赛级推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。