arXiv:2505.08364cs.AI2025-05EMNLP被引 27

模仿人类学习方式,让大模型更擅长解复杂数学题。

Learning Like Humans: Advancing LLM Reasoning Capabilities via Adaptive Difficulty Curriculum Learning and Expert-Guided Self-Reformulation

  • 动态调整训练难度,避免模型因题目变难而停滞
  • 引导模型重新表述专家解法,提升理解深度
  • 在AIME竞赛题上性能提升超16%,适合想提升推理能力的研究者

尽管大语言模型在数学推理等领域取得显著进展,但在稳定解决复杂问题方面仍面临挑战。受人类学习策略启发,本文提出两种新方法:自适应难度课程学习(ADCL),通过周期性重估后续数据批次的难度,缓解模型对题目难度感知随训练变化的问题;专家引导的自我重构(EGSR),一种强化学习策略,引导模型在自身概念框架内重构专家解法,而非直接模仿,从而促进深层理解与知识内化。在以Qwen2.5-7B为基础模型的多个挑战性数学推理基准上进行的大量实验表明,这两种策略协同作用显著提升性能。尤其在AIME24和AIME25基准上,联合使用相比标准零样本强化学习基线分别提升10%和16.6%。

原文摘要 · Abstract (English)

Despite impressive progress in areas like mathematical reasoning, large language models still face significant challenges in consistently solving complex problems. Drawing inspiration from key human learning strategies, we propose two novel strategies to enhance the capability of large language models to solve these complex problems. First, Adaptive Difficulty Curriculum Learning (ADCL) is a novel curriculum learning strategy that tackles the Difficulty Shift phenomenon (i.e., a model's perception of problem difficulty dynamically changes during training) by periodically re-estimating difficulty within upcoming data batches to maintain alignment with the model's evolving capabilities. Second, Expert-Guided Self-Reformulation (EGSR) is a novel reinforcement learning strategy that bridges the gap between imitation learning and pure exploration by guiding models to reformulate expert solutions within their own conceptual framework, rather than relying on direct imitation, fostering deeper understanding and knowledge assimilation. Extensive experiments on challenging mathematical reasoning benchmarks, using Qwen2.5-7B as the base model, demonstrate that these human-inspired strategies synergistically and significantly enhance performance. Notably, their combined application improves performance over the standard Zero-RL baseline by 10% on the AIME24 benchmark and 16.6% on AIME25.

大模型推理课程学习数学推理强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。