arXiv:2608.09538cs.CLcs.AI2026-08被引 1

评测大模型生成理论计算机科学证明的能力,构建首个研究级证明基准

TCS-BENCH: Benchmarking State-of-the-Art Generative AI Theoretical Computer Science Research Ability

  • 构建TCS-Bench,包含顶级会议论文中的定理证明任务
  • 验证器对专家标注结果的准确率超90%
  • 适合关注AI辅助数学证明的研究者

我们提出TCS-Bench,一个用于评估大语言模型(LLMs)在研究级理论计算机科学(TCS)证明生成能力的基准。TCS-Bench包含来自顶级理论计算机科学会议(STOC、FOCS、SODA)论文中的定理证明任务,每项任务提供推导目标结论所需的完整上下文。我们在该基准上评估了前沿模型,并通过验证代理确认生成证明的正确性。进一步将验证器与人类专家对一组目标命题与生成证明配对的判断进行对比。参考验证器在专家标注集上的准确率超过90%。

原文摘要 · Abstract (English)

We introduce TCS-Bench, a benchmark for evaluating Large Language Models (LLMs) on research-level Theoretical Computer Science (TCS) proof generation. TCS-Bench consists of theorem-proving tasks from papers published at top theoretical computer science venues (STOC, FOCS, and SODA). Each task provides the necessary context to derive a self-contained proof for a target result. We evaluate state-of-the-art models on this benchmark. We verify the correctness of generated proofs via a verification agent, and further benchmark the verifier against human-expert proof judgements on a set of target statements and generated proofs pairs. Our reference verifier achieves over 90% accuracy on the expert labeled set.

大模型证明生成理论计算机基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。