arXiv:2506.10481cs.AI2025-06被引 6

打造奥赛级编程难题库,评估大模型算法推理能力。

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics

  • 构建250道原创奥数级编程题,覆盖多种编程范式。
  • 顶尖模型在正确率和效率上已超多数人类参赛者。
  • 开源可测,适合研究代码生成与逻辑推理的学者。

随着模型日益复杂,传统算法评测数据集日趋饱和,亟需更具挑战性的基准来推动算法推理能力的发展。本文提出OIBench,一个高质量、私有且具有挑战性的奥数级信息学数据集,包含250道精心设计的原创题目。我们详述了基准的构建方法,确保在多种编程范式和复杂度上的全面评估,并通过实验验证其抗污染特性。提出时间/空间完成曲线,实现更精细的效率分析;并通过高层次参赛者评价,支持直接的人机对比。实验表明,尽管开源模型仍落后于闭源模型,但当前最优模型在正确性和效率上已超越大多数人类参赛者,但仍逊于标准解法。OIBench已作为完全开源资源发布(https://huggingface.co/datasets/AGI-Eval/OIBench),旨在推动未来大模型代码推理能力的进步。

原文摘要 · Abstract (English)

As models become increasingly sophisticated, conventional algorithm benchmarks are increasingly saturated, underscoring the need for more challenging benchmarks to guide future improvements in algorithmic reasoning. This paper introduces OIBench, a high-quality, private, and challenging olympiad-level informatics dataset comprising 250 carefully curated original problems. We detail the construction methodology of the benchmark, ensuring a comprehensive assessment across various programming paradigms and complexities, and we demonstrate its contamination-resistant properties via experiments. We propose Time/Space Completion Curves for finer-grained efficiency analysis and enable direct human-model comparisons through high-level participant evaluations. Our experiments reveal that while open-source models lag behind closed-source counterparts, current SOTA models already outperform most human participants in both correctness and efficiency, while still being suboptimal compared to the canonical solutions. By releasing OIBench as a fully open-source resource (https://huggingface.co/datasets/AGI-Eval/OIBench), we hope this benchmark will contribute to advancing code reasoning capabilities for future LLMs.

算法推理代码生成评测基准大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。