用形式化证明验证大模型数学推理步骤,杜绝幻觉。
Safe: Enhancing Mathematical Reasoning in Large Language Models via Retrospective Step-aware Formal Verification

- 每步推理转为Lean 4语言并形式化证明
- 在多个数据集上提升推理准确率
- 适合需要可验证推理的AI安全研究
链式思考(CoT)提示已成为激发大语言模型(LLM)推理能力的标准方法。然而,为缓解CoT中难以检测的幻觉问题,现有方法如过程奖励模型(PRMs)或自一致性机制如同黑箱,无法提供可核查证据,可能限制其效果。为此,我们借鉴“数学命题的黄金标准是提供证明”的理念,提出一种事后、步骤感知的形式化验证框架Safe。该框架不赋予任意评分,而是将每一步推理转化为形式化数学语言Lean 4,并提供形式化证明以识别幻觉。我们在多个语言模型和多种数学数据集上评估Safe,显著提升性能,并提供可解释、可验证的证据。我们还构建了FormalStep基准,包含30,809个形式化陈述,用于步骤正确性定理证明。据我们所知,这是首个将形式化数学语言Lean 4用于验证LLM生成自然语言内容的工作,契合形式化语言最初诞生的目的:为易出错的人类证明提供可靠基础。
原文摘要 · Abstract (English)
Chain-of-Thought (CoT) prompting has become the de facto method to elicit reasoning capabilities from large language models (LLMs). However, to mitigate hallucinations in CoT that are notoriously difficult to detect, current methods such as process reward models (PRMs) or self-consistency operate as opaque boxes and do not provide checkable evidence for their judgments, possibly limiting their effectiveness. To address this issue, we draw inspiration from the idea that "the gold standard for supporting a mathematical claim is to provide a proof". We propose a retrospective, step-aware formal verification framework $Safe$. Rather than assigning arbitrary scores, we strive to articulate mathematical claims in formal mathematical language Lean 4 at each reasoning step and provide formal proofs to identify hallucinations. We evaluate our framework $Safe$ across multiple language models and various mathematical datasets, demonstrating a significant performance improvement while offering interpretable and verifiable evidence. We also propose $FormalStep$ as a benchmark for step correctness theorem proving with $30,809$ formal statements. To the best of our knowledge, our work represents the first endeavor to utilize formal mathematical language Lean 4 for verifying natural language content generated by LLMs, aligning with the reason why formal mathematical languages were created in the first place: to provide a robust foundation for hallucination-prone human-written proofs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。