提出CREDIT方法,让自蒸馏奖励更聚焦输入特定推理。
From Generic Correlation to Input-Specific Credit in On-Policy Self Distillation

- 用贝叶斯滤波分析自蒸馏奖励,揭示其本质是点互信息
- 设计对比基线,分离出输入特定的奖励贡献
- 在代码、科学推理等任务上提升性能,几乎不增加计算开销
在线策略自蒸馏已成为后训练语言模型的有前景范式,模型通过环境反馈作为自身教师,提供无需外部教师或步骤标注的密集标记级奖励。尽管实证成功,但该奖励实际衡量什么以及如何分配信用仍不明确。在反馈条件的后验相容解释下,我们发现自蒸馏标记奖励是贝叶斯滤波增量,其轨迹和恰好等于给定输入时响应与反馈间的点互信息(pMI)。pMI可通过输入特定推理或通用捷径提升,因此我们沿输入轴分解教师对数概率。基于此,提出CREDIT(对比自蒸馏奖励),利用批对比基线分离输入特定成分。在序列层面,CREDIT是教师侧对对比pMI目标的代理,同时惩罚在无关输入下仍高概率的响应。在两种模型族上的编码、科学推理和工具使用基准测试中,CREDIT实现最强综合表现,额外计算开销可忽略。
原文摘要 · Abstract (English)
On-policy self-distillation has emerged as a promising paradigm for post-training language models, in which the model conditions on environment feedback to serve as its own teacher, providing dense token-level rewards without external teacher models or step-level annotations. Despite its empirical success, what this reward actually measures and what kind of credit it assigns remain unclear. Under a posterior-compatibility interpretation of feedback conditioning, standard in the implicit-reward literature, we show that the self-distillation token reward is a Bayesian filtering increment whose trajectory sum is exactly the pointwise mutual information between the response and the feedback given the input. This pMI can be raised by input-specific reasoning or by input-generic shortcuts, so we further decompose the teacher log-probability along the input axis. Based on this analysis, we propose CREDIT (Contrastive REward from DIsTillation), which isolates the input-specific component with a batch-contrastive baseline. At the sequence level, CREDIT is a teacher-side surrogate for a contrastive pMI objective that also penalizes responses remaining likely under unrelated inputs. Across coding, scientific reasoning, and tool-use benchmarks on two model families, CREDIT delivers the strongest aggregate performance at negligible additional compute.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。