评测大模型在理论计算机科学中的全流程研究能力,发现自动形式化是最大瓶颈。
FormalTCS: Benchmarking End-to-End Frontier Formal Theoretical Computer Science Research of Large Language Models

- 构建专家验证的端到端TCS研究基准,包含143个真实论文实例。
- 最佳模型形式化自然语言仅达11.5分,远低于人类提供形式化后的28.6分。
- 框架生成64条新命题,仅6条通过专家评估,体现研究判断力不足。
大语言模型在自动化理论计算机科学(TCS)研究中展现出潜力,但现有基准与真实研究场景差距较大。我们提出 ourbenchmark,一个专家验证的基准,用于评估大模型在前沿、端到端TCS研究中的表现。该基准包含143个来自STOC、FOCS、SODA和COLT 2025–2026年接收论文的实例,保留原始定义、假设和证明依赖关系,并配有专家验证的Lean形式化与证明。对领先LLMs的评估显示,当前模型尚未能可靠完成完整研究流程。其中,自动形式化是最大瓶颈:最佳模型在将自然语言命题转化为形式定理时仅得11.5分,而人类提供形式化后,证明任务达到28.6分(Pass@8)。基于此基准,我们进一步开发了自动化TCS研究框架,可生成、形式化、过滤并证明新命题。在64个生成命题中,仅有6个通过专家评估与证明验证,表明除形式化外,研究品味有限仍是自主研究的主要障碍。
原文摘要 · Abstract (English)
Large language models (LLMs) have shown growing potential for automated theoretical computer science (TCS) research, yet existing benchmarks remain far from realistic research settings. We introduce \ourbenchmark, an expert-validated benchmark for evaluating LLMs on frontier, end-to-end TCS research. \ourbenchmark contains $143$ instances drawn from papers accepted to STOC, FOCS, SODA, and COLT in 2025-2026, preserving paper-specific definitions, assumptions, and proof dependencies, with expert-verified Lean formalizations and proofs. Evaluations of leading LLMs reveal that current models remain far from reliably completing the full research pipeline. In particular, autoformalization is the sharpest bottleneck: the best model achieves only $11.5$ on translating natural-language claims into formal theorem statements, compared with $28.6$ Pass@8 when proving human-provided formal statements. Building on \ourbenchmark, we further develop an automated TCS research framework that generates, formalizes, filters, and proves new claims. Of $64$ generated claims, only $6$ ultimately pass expert evaluation and proof verification, indicating that beyond formalization, limited research taste remains another major barrier to autonomous TCS research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。