用逻辑谜题评估大模型反思纠错能力,提升数学推理表现。
FINEREASON: Evaluating and Improving LLMs' Deliberate Reasoning through Reflective Puzzle Solving
- 设计可拆解的逻辑谜题,细粒度检验推理中间步骤。
- 引入状态检查与转移任务,量化模型自我修正能力。
- 训练数据使数学推理准确率提升最高5.1%,适合模型优化研究者。
许多复杂推理任务不仅需要快速直觉反应,更需多步深思熟虑的策略。当前大语言模型(LLMs)正从‘系统1’快速响应转向‘系统2’反思修正型求解。然而现有评测仍过度依赖最终答案正确率,忽视了中间推理过程的审视,无法评估模型在推理中发现并纠正错误的能力。为此,我们提出FINEREASON,一个面向逻辑谜题的细粒度评测基准。每个谜题可分解为原子步骤,便于精确验证中间结果的正确性。在此基础上,我们构建两个任务:状态检查与状态转移,全面评估模型对当前情境的判断和下一步规划能力。为支持广泛研究,我们还提供一个用于训练的谜题数据集,旨在提升通用数学任务表现。实验表明,基于状态检查与转移数据训练的模型,在GSM8K数据集上数学推理准确率最高提升5.1%。
原文摘要 · Abstract (English)
Many challenging reasoning tasks require not just rapid, intuitive responses, but a more deliberate, multi-step approach. Recent progress in large language models (LLMs) highlights an important shift from the "System 1" way of quick reactions to the "System 2" style of reflection-and-correction problem solving. However, current benchmarks heavily rely on the final-answer accuracy, leaving much of a model's intermediate reasoning steps unexamined. This fails to assess the model's ability to reflect and rectify mistakes within the reasoning process. To bridge this gap, we introduce FINEREASON, a logic-puzzle benchmark for fine-grained evaluation of LLMs' reasoning capabilities. Each puzzle can be decomposed into atomic steps, making it ideal for rigorous validation of intermediate correctness. Building on this, we introduce two tasks: state checking, and state transition, for a comprehensive evaluation of how models assess the current situation and plan the next move. To support broader research, we also provide a puzzle training set aimed at enhancing performance on general mathematical tasks. We show that models trained on our state checking and transition data demonstrate gains in math reasoning by up to 5.1% on GSM8K.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。