arXiv:2505.18952cs.LG2025-05

用奖励引导强化学习,让小模型更好模仿大模型的偏好。

Online Knowledge Distillation with Reward Guidance

  • 设计奖励引导的模仿学习框架,动态优化学生与教师模型差异
  • 在偏好对齐中保持近最优,使学生模型性能逼近教师水平
  • 适用于在线和离线数据,支持白盒知识蒸馏场景

本文研究通过偏好优化进行大语言模型的知识蒸馏。提出一种基于奖励引导的序列式知识蒸馏框架,将学生与教师策略间的性能差距最小化,构建政策与奖励模型之间的极小极大优化问题。具体而言,奖励优化被约束在偏好对齐的置信集内以实现近似最优。针对偏好数据构建,探索了离线与在线两种知识蒸馏方式。此外,将奖励模型重构为$Q$-值函数形式,并将框架扩展至白盒知识蒸馏场景,此时教师模型的预测概率可访问。理论分析与实证结果均验证了该框架的有效性。

原文摘要 · Abstract (English)

This work studies knowledge distillation (KD) for large language models (LLMs) through preference optimization. We propose a reward-guided imitation learning framework for sequential KD, formulating a min-max optimization problem between the policy and reward model (RM) to minimize the performance gap between the student and teacher policies. Specifically, the reward optimization is constrained to achieve near-optimality within a confidence set for preference alignment. For preference data construction, we explore both offline and online preference-based KD. Additionally, we reformulate the RM using the $Q$-value function and extend the framework to white-box KD, where the teacher policy's predicted probabilities are accessible. Theoretical analysis and empirical results demonstrate the effectiveness of the proposed framework.

知识蒸馏偏好优化强化学习大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。