无监督强化学习可让大模型训练突破标注瓶颈,但有潜在崩溃风险。
How Far Can Unsupervised RLVR Scale LLM Training?
- 按奖励来源分内在与外在方法,内在法依赖模型自身信号
- 内在奖励随训练先升后降,崩溃时间由模型先验决定
- 提出新指标衡量模型先验,指导小数据测试训练
无监督强化学习结合可验证奖励(URLVR)为突破大语言模型训练的标注瓶颈提供了路径,通过无需真实标签的奖励机制实现扩展。现有方法利用模型内部信号,在早期展现良好效果,但其潜力与局限尚不明确。本文系统分析了URLVR,提出分类框架并建立统一理论:所有内在方法最终趋向于强化模型初始分布。当初始置信度与正确性一致时有效,反之则崩溃。实验表明,内在奖励普遍呈现先升后降趋势,崩溃时机由模型先验决定,而非工程设计。尽管存在扩展限制,内在奖励在小规模数据测试训练中仍具价值,并提出“模型坍塌步数”作为衡量模型先验的实用指标。此外,探索基于计算不对称性的外部奖励方法,初步显示可能突破置信度-正确性天花板。研究明确了内在URLVR的边界,也指明了可扩展替代方案的方向。
原文摘要 · Abstract (English)
Unsupervised reinforcement learning with verifiable rewards (URLVR) offers a pathway to scale LLM training beyond the supervision bottleneck by deriving rewards without ground truth labels. Recent works leverage model intrinsic signals, showing promising early gains, yet their potential and limitations remain unclear. In this work, we revisit URLVR and provide a comprehensive analysis spanning taxonomy, theory and extensive experiments. We first classify URLVR methods into intrinsic versus external based on reward sources, then establish a unified theoretical framework revealing that all intrinsic methods converge toward sharpening the model's initial distribution This sharpening mechanism succeeds when initial confidence aligns with correctness but fails catastrophically when misaligned. Through systematic experiments, we show intrinsic rewards consistently follow a rise-then-fall pattern across methods, with collapse timing determined by model prior rather than engineering choices. Despite these scaling limits, we find intrinsic rewards remain valuable in test-time training on small datasets, and propose Model Collapse Step to measure model prior, serving as a practical indicator for RL trainability. Finally, we explore external reward methods that ground verification in computational asymmetries, showing preliminary evidence they may escape the confidence-correctness ceiling. Our findings chart boundaries for intrinsic URLVR while motivating paths toward scalable alternatives.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。