arXiv:2507.11275cs.CL2025-07中稿 · ICML被引 5

用大模型自动把数学竞赛题转为形式化语言,构建了奥运级评测数据集。

FMC: Formalization of Natural Language Mathematical Competition Problems

  • 基于大模型与错误反馈的全自动无训练形式化流程
  • 建成3922道自然语言题与9787条Lean形式化对应数据,64.46%质量达标
  • 验证了少样本学习与迭代反馈能显著提升形式化效果,适合自动化推理研究

高效准确的自动形式化方法对于推进形式化数学推理至关重要。本文提出一种基于大语言模型并结合错误反馈的自动形式化流水线,实现完全自动且无需训练的形式化方法。利用该流程,我们构建了一个与奥林匹克级别数学问题对齐的Lean形式化数据集。该数据集包含3,922道自然语言数学题和9,787条Lean形式化表达,其中64.46%被评估为至少达到中等以上质量,适合作为自动化定理证明器的基准。此外,我们研究了多种大模型的形式化与推理能力,实证表明少样本学习、错误反馈及增加采样次数可有效提升自动形式化性能。在该数据集上对三种自动化定理证明器的实验也凸显了其挑战性,证实了其作为形式化推理任务基准的价值。

原文摘要 · Abstract (English)

Efficient and accurate autoformalization methods, which leverage large-scale datasets of extensive natural language mathematical problems to construct formal language datasets, are key to advancing formal mathematical reasoning. In this paper, we propose an autoformalization pipeline based on large language models with error feedback, achieving a fully automatic and training-free formalization approach. Using this pipeline, we curate an Olympiad-level dataset aligning natural language problems with Lean formalizations. The dataset comprises $3,922$ mathematical problems in natural language and $9,787$ in Lean, of which $64.46\%$ were assessed as at least above-average quality, making it suitable as a benchmark for automated theorem provers. Additionally, we investigate the formalization and reasoning capabilities of various LLMs and empirically demonstrate that few-shot learning, error feedback, and increasing sampling numbers enhance the autoformalization process. Experiments of three automated theorem provers on the \dataset\ dataset also highlight its challenging nature and its value as a benchmark for formal reasoning tasks.

形式化数学推理大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。