对比强化学习中任务奖励与分布锐化,发现前者更有效提升模型能力。
Beyond Distribution Sharpening: The Importance of Task Rewards
- 用强化学习实现分布锐化和任务奖励两种训练方式对比
- 分布锐化效果有限,且学习过程不稳定,最优解不佳
- 任务奖励能显著提升数学任务表现,适合追求稳定性能的场景
前沿模型在训练中引入基于任务奖励的强化学习后展现出卓越能力,使系统从纯推理模型进化为复杂智能体。然而,学界仍争论强化学习是否真正赋予模型新技能,还是仅通过分布锐化激发其潜在能力。为此,本文通过强化学习工具明确比较了分布锐化与任务奖励学习两种范式。分析表明,分布锐化存在根本性局限,从原理上揭示其最优解不理想且过程不稳。实验使用 Llama-3.2-3B-Instruct、Qwen2.5-3B-Instruct 及 Qwen3-4B-Instruct-2507 在数学数据集上的结果证实,分布锐化带来的提升有限,而引入任务奖励信号可显著实现稳健性能提升。
原文摘要 · Abstract (English)
Frontier models have demonstrated exceptional capabilities following the integration of task-reward-based reinforcement learning (RL) into their training pipelines, enabling systems to evolve from pure reasoning models into sophisticated agents. However, debate persists regarding whether RL genuinely instills new skills within a base model or merely sharpens its existing distribution to elicit latent capabilities. To address this dichotomy, we present an explicit comparison between distribution sharpening and task-reward-based learning, utilizing RL as a tool to implement both paradigms. Our analysis reveals the inherent limitations of distribution sharpening, demonstrating from first principles how and why the optima can be unfavorable and the approach fundamentally unstable. Furthermore, our experiments using Llama-3.2-3B-Instruct, Qwen2.5-3B-Instruct and Qwen3-4B-Instruct-2507 on math datasets confirm that sharpening yields limited gains, whereas incorporating task-based reward signal can greatly help achieve robust performance improvements and stable learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。