arXiv:2502.08908cs.AI2025-02

用强化学习优化大模型,让其更会自动证明定理。

Reinforced Large Language Model is a formal theorem prover

  • 用强化学习迭代优化预训练大模型的推理策略。
  • 在定理证明任务中准确率显著高于直接微调模型。
  • 适合对自动化证明和形式化验证感兴趣的读者。

为利用大语言模型在定理形式化和证明中的潜力,我们提出一种强化学习框架,通过滚动生成下一步策略并与其预期结果对比,迭代优化预训练的LLM。实验表明,该方法相较于直接微调的LLM能取得更高的准确率。

原文摘要 · Abstract (English)

To take advantage of Large Language Model in theorem formalization and proof, we propose a reinforcement learning framework to iteratively optimize the pretrained LLM by rolling out next tactics and comparing them with the expected ones. The experiment results show that it helps to achieve a higher accuracy compared with directly fine-tuned LLM.

定理证明强化学习大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。