arXiv:2506.04821cs.LG2025-06被引 2

用强化学习玩逻辑谜题,让大模型学会更可靠的数学推理。

LogicPuzzleRL: Cultivating Robust Mathematical Reasoning in LLMs via Reinforcement Learning

  • 通过七类自定义逻辑谜题进行强化学习,鼓励模型试错迭代。
  • 在中等难度的数学题上,泛化能力显著提升,尤其擅长多步推理。
  • 适合想提升模型通用推理能力的研究者或开发者使用。

大型语言模型在监督任务中表现优异,但在陌生场景下的结构化推理常显不足。这表明标准微调可能培养的是特定领域的启发式策略,而非通用思维模式。本文提出一种“边玩边学”框架,通过强化学习在七类定制逻辑谜题上微调大模型,每类谜题旨在培养约束传播、空间一致性、符号推理等不同能力。采用可验证奖励机制,模型根据解题正确性获得二元反馈,推动其进行假设驱动的迭代求解。实验显示,该训练方式显著提升了模型在多种数学基准上的分布外性能,尤其在需多步推理的中等难度问题上表现突出。分析表明,谜题训练促进了可迁移的推理模式,强化了代数运算、几何推断与组合逻辑能力,但在机械记忆或高度专业任务上提升有限。结果说明,基于逻辑谜题的强化学习重塑了大模型的内部推理机制,实现无需依赖特定符号工具的鲁棒且组合式泛化。

原文摘要 · Abstract (English)

Large language models (LLMs) excel at many supervised tasks but often struggle with structured reasoning in unfamiliar settings. This discrepancy suggests that standard fine-tuning pipelines may instill narrow, domain-specific heuristics rather than fostering general-purpose thinking strategies. In this work, we propose a "play to learn" framework that fine-tunes LLMs through reinforcement learning on a suite of seven custom logic puzzles, each designed to cultivate distinct reasoning skills such as constraint propagation, spatial consistency, and symbolic deduction. Using a reinforcement learning setup with verifiable rewards, models receive binary feedback based on puzzle correctness, encouraging iterative, hypothesis-driven problem solving. We demonstrate that this training approach significantly improves out-of-distribution performance on a range of mathematical benchmarks, especially for mid-difficulty problems that require multi-step reasoning. Analyses across problem categories and difficulty levels reveal that puzzle training promotes transferable reasoning routines, strengthening algebraic manipulation, geometric inference, and combinatorial logic, while offering limited gains on rote or highly specialized tasks. These findings show that reinforcement learning over logic puzzles reshapes the internal reasoning of LLMs, enabling more robust and compositional generalization without relying on task-specific symbolic tools.

大模型推理强化学习逻辑谜题数学能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。