arXiv:2504.11741cs.AIcs.CL2025-04被引 25

分析大模型数学推理能力,发现四层难度阶梯与突破瓶颈的关键。

Climbing the Ladder of Reasoning: What LLMs Can-and Still Can't-Solve after SFT?

  • 按难度分四层,不同层级需不同推理策略
  • 难题准确率卡在65%,对数据量依赖性强
  • 极端难题需非常规解法,当前模型普遍无法应对

近期的监督微调(SFT)方法显著提升了语言模型在数学推理任务上的表现,即使在小规模训练下亦然。然而,此类微调具体增强了哪些能力仍不清晰。本文通过在AIME24数据集上的细致分析,揭示了问题难度呈现阶梯状结构,将题目分为四个层级(易、中、难、极难)。研究发现,从易到中级需采用R1推理风格,仅需500-1000条微调样本即可实现跃升;而中级以上题目常因推理链中每一步出错导致准确率停滞在约65%,即便采用对数规模扩增数据也难以突破。极难题则带来根本性挑战,要求非传统解题技能,当前模型普遍无法处理。此外,精心构建的小规模数据集优势有限,扩大数据规模才更有效。该分析为提升语言模型数学推理能力提供了更清晰路径。

原文摘要 · Abstract (English)

Recent supervised fine-tuning (SFT) approaches have significantly improved language models' performance on mathematical reasoning tasks, even when models are trained at a small scale. However, the specific capabilities enhanced through such fine-tuning remain poorly understood. In this paper, we conduct a detailed analysis of model performance on the AIME24 dataset to understand how reasoning capabilities evolve. We discover a ladder-like structure in problem difficulty, categorize questions into four tiers (Easy, Medium, Hard, and Extremely Hard (Exh)), and identify the specific requirements for advancing between tiers. We find that progression from Easy to Medium tier requires adopting an R1 reasoning style with minimal SFT (500-1K instances), while Hard-level questions suffer from frequent model's errors at each step of the reasoning chain, with accuracy plateauing at around 65% despite logarithmic scaling. Exh-level questions present a fundamentally different challenge; they require unconventional problem-solving skills that current models uniformly struggle with. Additional findings reveal that carefully curated small-scale datasets offer limited advantage-scaling dataset size proves far more effective. Our analysis provides a clearer roadmap for advancing language model capabilities in mathematical reasoning.

数学推理大模型微调难度阶梯

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。