arXiv:2606.10479cs.AI2026-06

评测大模型在奥数组合数学中的推理与构造能力,发现两者差异明显。

ComBench: A Benchmark for Rigorous Proof Reasoning and Constructive Realization in Olympiad-Level Combinatorics

论文配图:ComBench: A Benchmark for Rigorous Proof Reasoning and Constructive Realization in Olympiad-Level Combinatorics
图 1 · 摘自论文原文
  • 构建两类奥数组合题:侧重证明的分析题和需构造解的实现实题。
  • 最强模型仅达65.4%准确率,构造题比存在性题更难。
  • 证明能力与构造能力可分离,适合评估模型创造性数学思维。

组合数学是奥数级数学解题的核心,需要深层离散推理、创造性构造与严谨结构洞察。近期证据表明,即使当前最强的大语言模型在奥数组合题上仍表现不均,暴露出创造性数学推理的差距。我们提出ComBench,一个面向奥数级组合数学的基准,用于评估与诊断大模型的组合推理能力。ComBench包含100道人工标注的竞赛级题目,分为两类互补场景:以分析为核心的题目(主要需严谨数学论证)和以构造为核心的题目(需显式构造并验证正确性)。评估协议结合评分标准指导的证明评分与确定性构造验证,揭示了证明质量与构造有效性之间的偏差。对前沿开源与闭源模型的实验显示,ComBench远未饱和:最强模型总体平均准确率达65.4%,最佳4次尝试中达到75.3%。进一步发现,严谨证明推理与构造实现是两种独立能力:Kimi-K2.6在分析型证明评分上落后于GPT-5.5,但在构造型问题的最佳4次尝试中超越后者;存在性问题与构造问题在主流前沿模型中始终最困难。

原文摘要 · Abstract (English)

Combinatorics is central to Olympiad-level mathematical problem solving, requiring deep discrete reasoning, creative constructions, and rigorous structural insight. Recent evidence suggests that even today's strongest frontier models remain uneven on Olympiad combinatorics, revealing a gap in creative mathematical reasoning. We introduce ComBench, an Olympiad-level combinatorics benchmark for evaluating and diagnosing the combinatorial reasoning capabilities of large language models. ComBench contains 100 human-annotated competition-level problems organized around two complementary settings: analysis-centric problems, which primarily require rigorous mathematical arguments, and construction-centric problems, which require explicit constructions in addition to correctness justifications. The evaluation protocol combines rubric-guided proof grading with deterministic construction verification, exposing cases where proof quality and construction validity diverge. Experiments on frontier open- and closed-source models show that ComBench is far from saturated: the strongest model reaches 65.4% overall Avg. and 75.3% overall Best@4. We further find that Rigorous Proof Reasoning and Constructive Realization are distinct capabilities: Kimi-K2.6 trails GPT-5.5 on analysis-centric proof grading but surpasses it on construction-centric Best@4, while Existence and Construction problems remain consistently hardest across representative frontier models.

组合数学大模型评测奥数构造能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。