用定理证明器自动评估并改进大模型生成的数学推理数据。
Theorem Prover as a Judge for Synthetic Data Generation
- 通过迭代自动形式化提升定理证明器的验证率,从60%到87%
- 利用定理证明器反馈替代人工标注,仅用3508样本提升模型性能
- 适合需要高质量数学推理数据的研究者和开发者
数学推理对合成数据的需求日益增长,因其有望提升大语言模型(LLM)的数学能力。然而,确保中间推理步骤的有效性仍是重大挑战,影响数据质量。尽管通过定理证明器进行形式化验证可有效检验LLM推理,但数学证明的自动形式化仍易出错。为此,我们提出迭代自动形式化,通过反复优化定理证明器形式化以减少错误,使在Lean证明器上的执行率从60%提升至87%。在此基础上,我们提出定理证明器作为裁判(TP-as-a-Judge),利用定理证明器形式化严格评估LLM中间推理,将自动形式化与合成数据生成深度融合。最后,我们提出基于定理证明器反馈的强化学习(RLTPF),以定理证明器反馈替代人类标注,取代人类反馈强化学习(RLHF)。在多个LLM上,应用TP-as-a-Judge与RLTPF,在仅3,508个样本下显著提升性能:在Mistral-7B上多步算术任务(MultiArith)提升5.56%,在Llama-2-7B上SVAMP任务提升6.00%,在Llama-3.1-8B上AQUA任务提升3.55%。
原文摘要 · Abstract (English)
The demand for synthetic data in mathematical reasoning has increased due to its potential to enhance the mathematical capabilities of large language models (LLMs). However, ensuring the validity of intermediate reasoning steps remains a significant challenge, affecting data quality. While formal verification via theorem provers effectively validates LLM reasoning, the autoformalisation of mathematical proofs remains error-prone. In response, we introduce iterative autoformalisation, an approach that iteratively refines theorem prover formalisation to mitigate errors, thereby increasing the execution rate on the Lean prover from 60% to 87%. Building upon that, we introduce Theorem Prover as a Judge (TP-as-a-Judge), a method that employs theorem prover formalisation to rigorously assess LLM intermediate reasoning, effectively integrating autoformalisation with synthetic data generation. Finally, we present Reinforcement Learning from Theorem Prover Feedback (RLTPF), a framework that replaces human annotation with theorem prover feedback in Reinforcement Learning from Human Feedback (RLHF). Across multiple LLMs, applying TP-as-a-Judge and RLTPF improves benchmarks with only 3,508 samples, achieving 5.56% accuracy gain on Mistral-7B for MultiArith, 6.00% on Llama-2-7B for SVAMP, and 3.55% on Llama-3.1-8B for AQUA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。