arXiv:2502.08908cs.AI2025-02
用强化学习优化大模型,让其更会自动证明定理。
Reinforced Large Language Model is a formal theorem prover
- 用强化学习迭代优化预训练大模型的推理策略。
- 在定理证明任务中准确率显著高于直接微调模型。
- 适合对自动化证明和形式化验证感兴趣的读者。
为利用大语言模型在定理形式化和证明中的潜力,我们提出一种强化学习框架,通过滚动生成下一步策略并与其预期结果对比,迭代优化预训练的LLM。实验表明,该方法相较于直接微调的LLM能取得更高的准确率。
原文摘要 · Abstract (English)
To take advantage of Large Language Model in theorem formalization and proof, we propose a reinforcement learning framework to iteratively optimize the pretrained LLM by rolling out next tactics and comparing them with the expected ones. The experiment results show that it helps to achieve a higher accuracy compared with directly fine-tuned LLM.
定理证明强化学习大模型
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。