arXiv:2606.22793cs.AI2026-06综述被引 2

系统梳理强化学习中策略蒸馏的反馈机制,提出可解释的理论框架。

A Formula-Driven Survey and Research Agenda for On-Policy Distillation

论文配图:A Formula-Driven Survey and Research Agenda for On-Policy Distillation
图 1 · 摘自论文原文
  • 从反馈到更新的统一公式视角,构建两类核心方法体系。
  • 揭示状态兼容性、词汇概率路由等关键影响因素,突破传统认知。
  • 适合对强化学习蒸馏机制感兴趣的研究者与工业应用开发者。

策略蒸馏(OPD)通过当前或近期学生策略生成状态序列,由教师或自教师在生成上下文中评分输出标记,将密集的概率、逻辑值或分布信号转化为后训练更新。本综述将OPD视为反馈到更新的问题,而非单一损失家族。我们基于两条路径——直接分布损失与策略梯度风格的对数比率更新——构建公式驱动的分类体系,并用于组织核心方法、验证器或结果引导的混合模型、工业报告、框架实现、失败模式与稳定化方案,均在明确证据边界下进行。该分类表明,OPD效果不仅取决于KL方向或教师访问,还受状态兼容性、支持构造、时间信用分配、词汇级概率路由、门控与权重设计及正则化的影响。我们进一步区分了常被混淆的两个机制:时间信用决定教师-学生对数比率回报如何加权回溯中的采样动作;词汇路由则决定当负面反馈抑制某采样标记时,概率质量应转向何处。这一区分确立了即时、未来回报、折扣与基线校正估计器的偏差边界,推动了基于价值的对数比率回报假设(GAE-OPD),并提出了反事实路由蒸馏(CR-OPD)以将概率质量导向教师支持且学生可达的替代路径。最后,我们将可行动诊断、失败机制、案例研究、开放问题与报告清单映射至同一反馈到更新变量空间。

原文摘要 · Abstract (English)

On-policy distillation (OPD) trains an LLM on states induced by the current or recent student policy: the student generates complete or partial rollouts, a teacher or self-teacher scores the resulting tokens under their generated contexts, and dense log-probability, logit, or distributional signals are converted into post-training updates. This survey studies OPD as a feedback-to-update problem rather than a single loss family. We develop a formula-driven taxonomy from two routes -- direct distributional losses and policy-gradient-style log-ratio updates -- and use it to organize core methods, verifier- or outcome-guided hybrids, industrial reports, framework implementations, failure modes, and stabilization recipes under explicit evidence boundaries. The taxonomy shows that OPD effectiveness depends not only on KL direction or teacher access, but also on state compatibility, support construction, temporal credit, vocabulary-level probability routing, gates and weights, and regularization. We further separate two mechanisms often conflated in sampled-token OPD stability discussions. Temporal credit asks how teacher-student log-ratio returns should weight sampled actions across a rollout; vocabulary routing asks where probability mass should move when negative feedback suppresses a sampled token. This distinction yields bias boundaries for immediate, return-to-go, discounted, and baseline-corrected estimators, motivates GAE-OPD as a value-based hypothesis for log-ratio returns, and motivates Counterfactual Routed OPD (CR-OPD) for routing probability mass toward teacher-supported, student-reachable alternatives. We close by mapping actionability diagnostics, failure mechanisms, case studies, open problems, and a reporting checklist onto the same feedback-to-update variables.

强化学习策略蒸馏反馈机制理论框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。