测试大模型在真实数学定义上的自动形式化能力,发现仍有明显短板。
Autoformalization in the Wild: Assessing LLMs on Real-World Mathematical Definitions
- 构建维基和arXiv真实数学定义数据集,评估大模型形式化能力。
- 模型自修正能力弱,未定义错误率高,与正式库对齐差。
- 引入外部反馈和上下文增强,可提升16%自纠错、43%减少未定义错误。
得益于语言能力,大模型有望通过自动形式化弥合非正式数学与形式语言之间的鸿沟。然而,其在复杂且自然出现的数学陈述上的泛化能力仍不明确。为此,我们研究了真实世界数学定义的自动形式化任务——数学话语的关键组成部分。具体地,我们引入两个新资源:从维基百科(Def_Wiki)和arXiv论文(Def_ArXiv)收集的定义。我们系统评估了多种大模型将定义形式化为Isabelle/HOL的能力,并探索了提升性能的策略,包括通过证明助手进行外部反馈优化,以及形式化定义的上下文锚定——即利用正式数学库中的相关元素增强大模型输出。结果表明,与miniF2F等现有基准相比,定义任务更具挑战性。大模型在自我修正和与正式库对齐方面仍存在显著困难。但结构化优化方法和定义锚定策略分别带来最高16%的自我修正提升和43%的未定义错误减少,凸显了在真实场景中增强大模型自动形式化能力的可行路径。
原文摘要 · Abstract (English)
Thanks to their linguistic capabilities, LLMs offer an opportunity to bridge the gap between informal mathematics and formal languages through autoformalization. However, it is still unclear how well LLMs generalize to sophisticated and naturally occurring mathematical statements. To address this gap, we investigate the task of autoformalizing real-world mathematical definitions: a critical component of mathematical discourse. Specifically, we introduce two novel resources for autoformalization, collecting definitions from Wikipedia (Def_Wiki) and arXiv papers (Def_ArXiv). We then systematically evaluate a range of LLMs, analyzing their ability to formalize definitions into Isabelle/HOL. Furthermore, we investigate strategies to enhance LLMs' performance including refinement through external feedback from Proof Assistants, and formal definition grounding, where we augment LLMs' formalizations through relevant contextual elements from formal mathematical libraries. Our findings reveal that definitions present a greater challenge compared to existing benchmarks, such as miniF2F. In particular, we found that LLMs still struggle with self-correction, and aligning with relevant mathematical libraries. At the same time, structured refinement methods and definition grounding strategies yield notable improvements of up to 16% on self-correction capabilities and 43% on the reduction of undefined errors, highlighting promising directions for enhancing LLM-based autoformalization in real-world scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。