用大模型生成可验证的正式证明,提升代码与数学推理能力
From Informal to Formal -- Incorporating and Evaluating LLMs on Natural Language Requirements to Verifiable Formal Proofs
- 构建1.8万条指令-响应对,覆盖5种形式化语言
- 小模型微调后达到6710亿参数大模型水平
- 微调后模型在数学、编程等任务上均有提升
基于AI的形式化数学推理研究呈现迅猛增长趋势,在国际数学奥林匹克竞赛等场景表现优异。本文聚焦形式化验证这一直接应用场景,将其分解为多个子任务。通过提炼GPT-4o生成内容,构建了涵盖五种形式化规范语言(Coq、Lean4、Dafny、ACSL和TLA+)的18,000条高质量指令-响应对,并在十款开源大模型(包括近期流行的DeepSeek-R1)上进行评估。同时,对多个7~8B规模的小模型进行微调,使其性能达到DeepSeek-R1-671B水平。有趣的是,使用形式化数据微调后,模型在数学、推理和编码能力方面均得到增强。相关微调模型已发布于https://huggingface.co/fm-universe。
原文摘要 · Abstract (English)
The research in AI-based formal mathematical reasoning has shown an unstoppable growth trend. These studies have excelled in mathematical competitions like IMO and have made significant progress. This paper focuses on formal verification, an immediate application scenario of formal reasoning, and breaks it down into sub-tasks. We constructed 18k high-quality instruction-response pairs across five formal specification languages (Coq, Lean4, Dafny, ACSL, and TLA+) by distilling gpt-4o and evaluated against ten open-sourced LLMs, including recent popular DeepSeek-R1. We also fine-tuned several 7~8B small models to achieve comparable performance with Deepseek-R1-671B. Interestingly, we observed that fine-tuning with formal data also enhances mathematics, reasoning, and coding capabilities. Fine-tuned models are released at https: //huggingface.co/fm-universe.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。