arXiv:2602.08222cs.AI2026-02被引 6

用模型过去的弱状态指导优化,突破大模型训练饱和瓶颈。

Weak-Driven Learning: How Weak Agents make Strong Agents Stronger

  • 通过熵动态识别模型历史弱态中的可恢复学习缺口
  • 在数学推理和代码生成任务上实现显著性能提升
  • 无需额外推理开销,适合高效迭代大模型

随着后训练优化成为提升大语言模型的核心手段,我们观察到一个持续存在的饱和瓶颈:当模型变得高度自信后,进一步训练带来的收益急剧下降。现有方法仍聚焦于强化目标预测,但我们发现模型自身历史弱态中仍蕴含着有信息量的监督信号。受此启发,我们提出WMSS(Weak Agents Can Make Strong Agents Stronger)——一种利用弱检查点引导持续优化的后训练范式。通过熵动态识别可恢复的学习差距,并以补偿性学习加以强化,WMSS使强模型突破传统后训练的饱和限制。在数学推理和代码生成数据集上的实验表明,采用该方法训练的智能体实现了有效性能提升,且不增加任何推理成本。

原文摘要 · Abstract (English)

As post-training optimization becomes central to improving large language models, we observe a persistent saturation bottleneck: once models grow highly confident, further training yields diminishing returns. While existing methods continue to reinforce target predictions, we find that informative supervision signals remain latent in models' own historical weak states. Motivated by this observation, we propose WMSS (Weak Agents Can Make Strong Agents Stronger), a post-training paradigm that leverages weak checkpoints to guide continued optimization. By identifying recoverable learning gaps via entropy dynamics and reinforcing them through compensatory learning, WMSS enables strong agents to improve beyond conventional post-training saturation. Experiments on mathematical reasoning and code generation datasets show that agents trained with our approach achieve effective performance improvements, while incurring zero additional inference cost.

大模型优化后训练弱监督性能提升

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。