arXiv:2508.04440cs.CLcs.AI2025-08AAAI被引 17

用知识与推理融合提升大模型自动形式化能力

StepFun-Formalizer: Unlocking the Autoformalization Potential of LLMs through Knowledge-Reasoning Fusion

  • 构建双数据集,融合形式化知识与推理路径
  • 32B模型在两个基准上达40.5%和26.7%准确率
  • 适合数学形式化、AI辅助证明等研究者使用

自动形式化旨在将自然语言数学陈述转化为形式语言。尽管大模型加速了该领域进展,现有方法仍存在准确率低的问题。我们识别出有效自动形式化的两大关键能力:对形式语言领域知识的全面掌握,以及对自然语言问题的理解与非形式-形式映射的推理能力。缺乏前者,模型无法识别正确形式对象;缺乏后者,难以精准理解现实语境并映射为形式表达。为此,我们提出ThinkingF数据合成与训练流程,以同时提升这两项能力。首先,通过提炼和筛选大规模富含形式知识的样本构建数据集;其次,基于专家设计模板生成从非形式到形式的推理轨迹。随后,利用SFT与RLVR在上述数据上训练,进一步融合与优化两种能力。最终得到的7B与32B模型兼具全面的形式知识与强大的非形式-形式推理能力。值得注意的是,StepFun-Formalizer-32B在FormalMATH-Lite上达到40.5%的BEq@1分数,在ProverBench上达到26.7%,超越所有先前通用与专用模型。

原文摘要 · Abstract (English)

Autoformalization aims to translate natural-language mathematical statements into a formal language. While LLMs have accelerated progress in this area, existing methods still suffer from low accuracy. We identify two key abilities for effective autoformalization: comprehensive mastery of formal-language domain knowledge, and reasoning capability of natural language problem understanding and informal-formal alignment. Without the former, a model cannot identify the correct formal objects; without the latter, it struggles to interpret real-world contexts and map them precisely into formal expressions. To address these gaps, we introduce ThinkingF, a data synthesis and training pipeline that improves both abilities. First, we construct two datasets: one by distilling and selecting large-scale examples rich in formal knowledge, and another by generating informal-to-formal reasoning trajectories guided by expert-designed templates. We then apply SFT and RLVR with these datasets to further fuse and refine the two abilities. The resulting 7B and 32B models exhibit both comprehensive formal knowledge and strong informal-to-formal reasoning. Notably, StepFun-Formalizer-32B achieves SOTA BEq@1 scores of 40.5% on FormalMATH-Lite and 26.7% on ProverBench, surpassing all prior general-purpose and specialized models.

形式化大模型数学推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。