用大模型把数学口语变严格证明,提升AI的可信推理能力
Autoformalization in the Era of Large Language Models: A Survey
- 用大模型将非正式数学命题转为可验证形式化表达
- 揭示了大模型生成内容的可验证性提升路径
- 适合关注AI可信推理与自动定理证明的研究者
自动形式化(autoformalization)是将非正式数学命题转化为可验证形式化表示的基础任务,在自动化定理证明中具有重要意义,为数学在理论与应用领域的使用提供了新视角。随着人工智能,特别是大语言模型(LLMs)的快速发展,该领域取得显著进展,带来新机遇与独特挑战。本文从数学与大模型双重视角,全面综述近期进展,涵盖不同数学领域与难度下的应用,分析从数据预处理到模型设计与评估的端到端流程。进一步探讨自动形式化在增强大模型输出可验证性中的新兴作用,凸显其提升大模型可信度与推理能力的潜力。最后,总结当前支持研究的关键开源模型与数据集,讨论开放挑战与未来方向。
原文摘要 · Abstract (English)
Autoformalization, the process of transforming informal mathematical propositions into verifiable formal representations, is a foundational task in automated theorem proving, offering a new perspective on the use of mathematics in both theoretical and applied domains. Driven by the rapid progress in artificial intelligence, particularly large language models (LLMs), this field has witnessed substantial growth, bringing both new opportunities and unique challenges. In this survey, we provide a comprehensive overview of recent advances in autoformalization from both mathematical and LLM-centric perspectives. We examine how autoformalization is applied across various mathematical domains and levels of difficulty, and analyze the end-to-end workflow from data preprocessing to model design and evaluation. We further explore the emerging role of autoformalization in enhancing the verifiability of LLM-generated outputs, highlighting its potential to improve both the trustworthiness and reasoning capabilities of LLMs. Finally, we summarize key open-source models and datasets supporting current research, and discuss open challenges and promising future directions for the field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。