arXiv:2512.15699cs.LGcs.SE2025-12被引 16

构建计算机科学前沿难题基准,评估模型解决未知最优解问题的能力。

FrontierCS: Evolving Challenges for Evolving Intelligence

  • 设计156个开放性问题,由专家评审并提供可执行代码评估方式。
  • 模型需编写可运行程序而非直接答答案,算法与科研类问题均含客观评分。
  • 当前模型仍远落后于人类,且过度优化可运行代码而非高质量算法设计。

我们提出FrontierCS,一个包含156个跨领域计算机科学开放性问题的基准,由计算机科学博士及顶级编程竞赛参与者和出题人共同设计与评审。不同于现有聚焦已知最优解的任务基准,FrontierCS针对最优解未知但解质量可客观评估的问题。模型通过实现可执行程序来解决问题,而非输出直接答案。该基准涵盖算法类问题(如竞赛题的NP难变体,支持部分评分)和科研类问题,每道题均提供专家参考解法与自动评测器。结合开放设计、可度量进展与专家筛选,FrontierCS构成计算机科学难度前沿的评测标准。实证表明,当前前沿推理模型在算法与科研任务上仍显著落后于人类专家,仅增加推理预算无法缩小差距,且模型常过度优化生成可运行代码,而非发现高质量算法与系统设计。

原文摘要 · Abstract (English)

We introduce FrontierCS, a benchmark of 156 open-ended problems across diverse areas of computer science, designed and reviewed by experts, including CS PhDs and top-tier competitive programming participants and problem setters. Unlike existing benchmarks that focus on tasks with known optimal solutions, FrontierCS targets problems where the optimal solution is unknown, but the quality of a solution can be objectively evaluated. Models solve these tasks by implementing executable programs rather than outputting a direct answer. FrontierCS includes algorithmic problems, which are often NP-hard variants of competitive programming problems with objective partial scoring, and research problems with the same property. For each problem we provide an expert reference solution and an automatic evaluator. Combining open-ended design, measurable progress, and expert curation, FrontierCS provides a benchmark at the frontier of computer-science difficulty. Empirically, we find that frontier reasoning models still lag far behind human experts on both the algorithmic and research tracks, that increasing reasoning budgets alone does not close this gap, and that models often over-optimize for generating merely workable code instead of discovering high-quality algorithms and system designs.

基准测试算法挑战智能推理开放问题

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。