arXiv:2505.23703cs.AIcs.CL2025-05EMNLP被引 12

将自然语言与形式化语言结合,显著提升大模型数学解题能力

Let's Reason Formally: Natural-Formal Hybrid Reasoning Enhances LLM's Math Capability

  • 用形式化语言重构自然语言数学题,实现双模态协同推理
  • 在MATH-500和AMC测试中准确率分别达89.80%和84.34%
  • 适合需要高精度数学推理的AI研究者与教育技术开发者

提升大模型的数学推理能力已成为数学与计算机领域的重要课题。尽管近期研究在纯自然语言(NL)与形式语言(FL)推理上取得进展,但强化学习方法难以赋予基模型未包含的新能力,亟需有效融合形式化知识。然而,NL与FL在问题结构与推理格式上的差异带来挑战。为此,我们提出端到端的NL-FL混合推理框架(NFL-HR),通过NL-FL问题对齐方法将自然语言问答题转化为形式语言中的存在性定理,并设计混合输入机制使形式语言推理器同时处理两类问题。此外,基于LLM的答案提取机制缓解了输出格式差异。实验表明,NFL-HR在MATH-500和AMC基准上分别达到89.80%和84.34%的准确率,较自然语言基线分别提升4.60%和4.82%。部分问题即使在更多尝试下,基线模型仍无法解决。

原文摘要 · Abstract (English)

Enhancing the mathematical reasoning capabilities of LLMs has garnered significant attention in both the mathematical and computer science communities. Recent works have made substantial progress in both Natural Language (NL) reasoning and Formal Language (FL) reasoning by leveraging the potential of pure Reinforcement Learning (RL) methods on base models. However, RL approaches struggle to impart new capabilities not presented in the base model, highlighting the need to integrate more knowledge like FL into NL math reasoning effectively. Yet, this integration is challenging due to inherent disparities in problem structure and reasoning format between NL and FL. To address these challenges, we introduce **NL-FL HybridReasoning (NFL-HR)**, an end-to-end framework designed to incorporate the FL expert into NL math problem-solving. To bridge the NL and FL input format gap, we propose the NL-FL Problem Alignment method, which reformulates the Question-Answering (QA) problems in NL as existence theorems in FL. Subsequently, the Mixed Problem Input technique we provide enables the FL reasoner to handle both QA and existence problems concurrently. Lastly, we mitigate the NL and FL output format gap in reasoning through an LLM-based Answer Extraction mechanism. Comprehensive experiments demonstrate that the NFL-HR framework achieves **89.80**% and **84.34%** accuracy rates on the MATH-500 and the AMC benchmarks, surpassing the NL baseline by **4.60%** and **4.82%**, respectively. Notably, some problems resolved by our framework remain unsolved by the NL baseline model even under a larger number of trials.

数学推理混合推理形式化语言大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。