用形式化验证实时纠错,让大模型推理更严谨
Pushing the Boundaries of Natural Reasoning: Interleaved Bonus from Formal-Logic Verification
- 动态穿插逻辑验证与生成过程,实时发现并修正错误
- 7B和14B模型在6个基准上分别超越基线10.4%和14.2%
- 适合需要高可靠推理的数学、逻辑类任务
大型语言模型虽具强大能力,但其基于概率的逐词生成易导致逻辑矛盾与奖励劫持,而形式符号系统可避免此类问题。为弥合这一差距,本文提出一种形式逻辑验证引导的框架,将形式化符号验证动态穿插于自然语言生成过程中,实现实时反馈以检测并修正错误。区别于以往仅被动事后验证的神经符号方法,本方法在推理链中主动惩罚中间谬误。通过新颖的两阶段训练流程,融合形式逻辑验证引导的监督微调与策略优化,实现高效协同。在涵盖数学、逻辑与通用推理的六个基准上进行的广泛评估表明,7B与14B模型平均性能分别优于当前最优基线10.4%与14.2%。结果验证了形式化验证可作为可扩展机制,显著提升先进大模型推理性能边界。
原文摘要 · Abstract (English)
Large Language Models (LLMs) show remarkable capabilities, yet their stochastic next-token prediction creates logical inconsistencies and reward hacking that formal symbolic systems avoid. To bridge this gap, we introduce a formal logic verification-guided framework that dynamically interleaves formal symbolic verification with the natural language generation process, providing real-time feedback to detect and rectify errors as they occur. Distinguished from previous neuro-symbolic methods limited by passive post-hoc validation, our approach actively penalizes intermediate fallacies during the reasoning chain. We operationalize this framework via a novel two-stage training pipeline that synergizes formal logic verification-guided supervised fine-tuning and policy optimization. Extensive evaluation on six benchmarks spanning mathematical, logical, and general reasoning demonstrates that our 7B and 14B models outperform state-of-the-art baselines by average margins of 10.4% and 14.2%, respectively. These results validate that formal verification can serve as a scalable mechanism to significantly push the performance boundaries of advanced LLM reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。