arXiv:2602.10014cs.LGstat.ML2026-02

提出任务导向的渐进式自提升理论,解释模型如何通过难易递进任务持续优化

A Task-Centric Theory for Iterative Self-Improvement with Easy-to-Hard Curricula

  • 将每轮自提升建模为奖励过滤数据上的最大似然微调
  • 证明易到难课程可使模型在相同样本预算下获得更优期望奖励
  • 揭示模型越强接收数据越多的正反馈机制,解释性能饱和原因

迭代自提升通过用语言模型自身生成并经过奖励验证的输出来微调该模型。尽管实践效果显著,但这种生成性、迭代过程在有限样本条件下的理论基础仍不充分。本文将每轮自提升建模为在奖励过滤分布上的最大似然微调,并推导出期望奖励的有限样本保证。分析揭示了明确的反馈循环:模型越优,每轮可接受的数据越多,支持持续自提升,同时解释了最终性能饱和的原因。采用任务中心视角,考虑具有多个难度级别的推理任务,进一步证明在模型初始化、任务难度与样本预算满足特定量化条件下,易到难课程优于固定混合任务训练。分析结果经蒙特卡洛模拟及合成图推理任务和多个标准数学推理基准实验验证。

原文摘要 · Abstract (English)

Iterative self-improvement fine-tunes an autoregressive large language model (LLM) on reward-verified outputs generated by the LLM itself. In contrast to the empirical success of self-improvement, the theoretical foundation of this generative, iterative procedure in a practical, finite-sample setting remains limited. We make progress toward this goal by modeling each round of self-improvement as maximum-likelihood fine-tuning on a reward-filtered distribution and deriving finite-sample guarantees for the expected reward. Our analysis reveals an explicit feedback loop where better models accept more data per iteration, supporting sustained self-improvement while explaining eventual saturation of such improvement. Adopting a task-centric view by considering reasoning tasks with multiple difficulty levels, we further prove quantifiable conditions on model initialization, task difficulty, and sample budget where easy-to-hard curricula provably achieve better guarantees than training on fixed mixtures of tasks. Our analyses are validated through Monte-Carlo simulations and experiments spanning a synthetic graph-based reasoning task and multiple standard mathematical reasoning benchmarks.

自提升课程学习大模型理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。