arXiv:2512.24796cs.LOcs.AI2025-12被引 4

构建首个形式化范畴论基准,测试AI在抽象结构推理上的真实能力

LeanCat: A Benchmark Suite for Formal Category Theory in Lean (Part I: 1-Categories)

  • 用100个形式化范畴论任务构建新基准,聚焦高阶接口推理
  • 顶尖模型仅12.0%通过率,难题全失败,暴露组合泛化缺陷
  • 引入检索增强代理实现动态纠错,性能翻倍,验证迭代必要性

尽管大型语言模型在形式化定理证明中展现出强大能力,但现有基准未能有效衡量基于库的抽象能力——即处理现代数学与软件工程中高阶接口和可复用结构的能力。我们提出LeanCat,一个包含100个在Lean中完全形式化的范畴论任务的挑战性基准。与代数或算术不同,范畴论对结构化、接口级推理构成严格压力测试。评估显示存在严重抽象鸿沟:最佳现有时态模型在pass@4下仅能解决12.0%的任务,且性能从易题的55.0%骤降至难题的0.0%,凸显组合泛化失败。为克服此问题,我们评估了LeanBridge,一种采用检索-生成-验证循环的检索增强代理。LeanBridge达到最高24.0%的成功率,是最佳静态基线的两倍。结果表明,在抽象领域中,迭代精炼与动态库检索并非优化手段,而是神经符号推理的刚性需求。LeanCat提供了一个紧凑、可复用的测试平台,用于追踪研究级形式化进展。

原文摘要 · Abstract (English)

While large language models (LLMs) have demonstrated impressive capabilities in formal theorem proving, current benchmarks fail to adequately measure library-grounded abstraction -- the ability to reason with high-level interfaces and reusable structures central to modern mathematics and software engineering. We introduce LeanCat, a challenging benchmark comprising 100 fully formalized category-theory tasks in Lean. Unlike algebra or arithmetic, category theory serves as a rigorous stress test for structural, interface-level reasoning. Our evaluation reveals a severe abstraction gap: the best state-of-the-art model solves only 12.0% of tasks at pass@4, with performance collapsing from 55.0% on Easy tasks to 0.0% on High-difficulty tasks, highlighting a failure in compositional generalization. To overcome this, we evaluate LeanBridge, a retrieval-augmented agent that employs a retrieve-generate-verify loop. LeanBridge achieves a peak success rate of 24.0% -- doubling the performance of the best static baseline. These results empirically demonstrate that iterative refinement and dynamic library retrieval are not merely optimizations but strict necessities for neuro-symbolic reasoning in abstract domains. LeanCat offers a compact, reusable testbed for tracking progress toward reliable, research-level formalization.

形式化证明范畴论神经符号推理基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。