arXiv:2505.20415cs.CL2025-05EMNLP被引 9

用蒙特卡洛方法生成符号化推理轨迹,提升大模型逻辑推理能力。

Enhancing Logical Reasoning in Language Models via Symbolically-Guided Monte Carlo Process Supervision

  • 通过蒙特卡洛估计合成带步骤伪标签的符号化推理路径。
  • 在FOLIO和LogicAsker上显著提升模型逻辑推理与泛化性能。
  • 适合关注推理可靠性与符号化规划的研究者或应用开发者。

大语言模型在多个推理基准上表现优异,但近期研究指出其性能主要源于记忆而非泛化能力。模型对内容变化敏感,缺乏稳健的规划或符号抽象支持。为提高可靠性,已有研究尝试结合符号方法,但受限于难以构建可靠且可扩展的验证机制。本文提出通过蒙特卡洛估计大规模合成高质量符号推理轨迹及逐步伪标签,进而训练过程奖励模型(PRM)以筛选更符号化的推理路径。这些路径随后用于直接偏好优化(DPO)与监督微调(SFT),提升逻辑推理与泛化能力。实验在FOLIO和LogicAsker基准上验证了方法有效性,对前沿与开源模型均有增益。额外在声明验证数据上的实验表明,基于生成符号轨迹的微调增强了跨领域泛化能力,显示出该方法在提升规划与逻辑推理方面的潜力。

原文摘要 · Abstract (English)

Large language models (LLMs) have shown strong performance in many reasoning benchmarks. However, recent studies have pointed to memorization, rather than generalization, as one of the leading causes for such performance. LLMs, in fact, are susceptible to content variations, demonstrating a lack of robust planning or symbolic abstractions supporting their reasoning process. To improve reliability, many attempts have been made to combine LLMs with symbolic methods. Nevertheless, existing approaches fail to effectively leverage symbolic representations due to the challenges involved in developing reliable and scalable verification mechanisms. In this paper, we propose to overcome such limitations by synthesizing high-quality symbolic reasoning trajectories with stepwise pseudo-labels at scale via Monte Carlo estimation. A Process Reward Model (PRM) can be efficiently trained based on the synthesized data and then used to select more symbolic trajectories. The trajectories are then employed with Direct Preference Optimization (DPO) and Supervised Fine-Tuning (SFT) to improve logical reasoning and generalization. Our results on benchmarks (i.e., FOLIO and LogicAsker) show the effectiveness of the proposed method with gains on frontier and open-weight models. Moreover, additional experiments on claim verification data reveal that fine-tuning on the generated symbolic reasoning trajectories enhances out-of-domain generalizability, suggesting the potential impact of the proposed method in enhancing planning and logical reasoning.

逻辑推理符号化强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。