arXiv:2604.06298cs.LG2026-04中稿 · ICLR

难样本对小模型推理提升有限,低难度数据训练更高效

Limits of Difficulty Scaling: Hard Samples Yield Diminishing Returns in GRPO-Tuned SLMs

  • 用GRPO+LoRA在30亿参数以下模型上优化数学推理
  • 难题训练收益递减,低难度数据仅用45%步数达同等效果
  • 跨数据集迁移时GSM8K训练反而在MATH上表现更好

近期大模型对齐研究认为偏好优化可通过转移概率质量提升推理能力。我们在资源受限场景下,采用GRPO与LoRA对最大30亿参数的轻量语言模型(SLMs)进行数学推理训练,覆盖GSM8K和MATH数据集,并按难度分层分析。随着题目难度上升,准确率趋于饱和,揭示出能力边界:GRPO主要改变输出偏好,无法可靠提升最难题目的解题能力。一致地,仅在低难度问题上训练GRPO,即可在各难度层级达到全数据集训练的准确率,且仅需约45%的训练步骤,表明在该范式下高难度样本回报递减。此外发现跨数据集泛化现象:在GSM8K上训练的GRPO,在MATH的数值子集上表现优于在MATH上训练的模型,分别领先约5%(1.5B)和3%(3B)。我们证明,最佳可实现增益高度依赖于基础模型的先验推理能力及数据集难度分布。

原文摘要 · Abstract (English)

Recent alignment work on Large Language Models (LLMs) suggests preference optimization can improve reasoning by shifting probability mass toward better solutions. We test this claim in a resource-constrained setting by applying GRPO with LoRA to SLMs (up to 3B) for math reasoning on GSM8K and MATH datasets with difficulty-stratified analyses. As problem difficulty increases, accuracy plateaus, revealing a capacity boundary: GRPO primarily reshapes output preferences without reliably improving hardest-tier solving. Consistent with this, training GRPO only on lower-difficulty problems matches full-dataset accuracy across difficulty tiers while using only ~45% training steps, indicating diminishing returns from harder samples in this regime. We also find a cross-dataset generalization effect: GSM8K-trained GRPO achieves higher accuracy on the numeric subset of MATH than MATH-trained GRPO, exceeding it by ~5% at 1.5B and by ~3% at 3B. We show that the best achievable gains depend strongly on the base model's prior reasoning competence and the dataset's difficulty profile.

推理增强小模型偏好优化数学推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。