用特权信息给令牌打分未必有效,需三重验证。
Privileged Likelihood Is Not Automatically Value: Three Checks for Token Credit in On-Policy Self-Distillation
- 通过三种检验评估令牌得分是否真实反映任务成功
- AIME2025上得分仅随机水平(AUC=0.505),最优模型仅33.9%准确率
- 仅当得分与结果高度相关时,自蒸馏才有效,否则性能远低于直接奖励
在策略自蒸馏旨在通过引用参考解或评判反馈等特权信息,提供令牌级评分来改进可验证奖励强化学习(RLVR)。这些评分被当作令牌级动作值的估计,但实质回答的是:输入上下文丰富后模型预测如何变化,而非改变一个令牌后预期结果如何变化。本文从三个维度检验这一差距:(i) 令牌级评分是否追踪任务成功;(ii) 同次推演生成的反馈是否导致评分自我参照;不同推演的反馈能否缓解此问题;(iii) 实际训练目标强化了何种行为。在AIME 2025上的实验显示,评分仅以接近随机水平区分正确与错误推演(AUC=0.505);使用其他推演反馈未显著提升判别力;所有配置的平均准确率仅为24.2–33.9% Avg@4,远低于仅基于结果的GRPO方法(64.2%)。高熵令牌十等分占总绝对优势质量的57–71%,但其评分对推理质量最不敏感。相反,在SciKnowEval Biology上,对应轨迹评分的AUC达0.81–0.92,使保持分数提升28.0%。结果表明,仅当基于似然的评分经实证验证为结果相关代理时,密集信用分配才有效;否则监督信号可能失效且严重劣于基于结果的强化学习。
原文摘要 · Abstract (English)
On-policy self-distillation aims to improve upon reinforcement learning from verifiable rewards (RLVR) by providing token-level scores derived from privileged information, such as reference solutions or critic feedback. These scores are treated as estimates of token-level action values, yet they answer a fundamentally different question: how the model's prediction changes when its input context is enriched, rather than how the expected outcome changes when a token is changed. We examine this gap along three dimensions: (i) whether the token-level score tracks task success; (ii) whether feedback generated from the same rollout causes the score to reflect agreement with its own description, and whether using feedback from other rollouts in the group mitigates this self-referential effect; and (iii) what behavior the resulting training objective actually reinforces. In experiments on AIME 2025, the implemented score distinguishes correct from incorrect rollouts at approximately chance level (AUC=0.505); using feedback from a different rollout does not consistently improve this discrimination; and all training configurations achieve only 24.2-33.9% Avg@4, compared with 64.2% for outcome-only GRPO. Moreover, the highest-entropy token decile accounts for 57-71% of the total absolute token-advantage mass, despite the score being least informative about reasoning quality in this regime. By contrast, similar experiments on SciKnowEval Biology improves held-out Avg@8 by 28.0%, while its corresponding trajectory scores achieve AUCs of 0.81-0.92. Together, these results suggest that dense credit assignment through distillation can be effective when its likelihood-based scores are empirically validated as meaningful proxies for outcome-relevant credit. When this alignment does not hold, however, the resulting supervision can fail to generalize and may substantially underperform outcome-based RL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。