arXiv:2504.19024cs.CL2025-04Conference of the …被引 1

用多步回报估计提升大模型文本生成的强化学习知识蒸馏效果

KETCHUP: K-Step Return Estimation for Sequential Knowledge Distillation

  • 基于贝尔曼最优方程设计多步回报,降低梯度估计方差
  • 在三个文本生成任务中均超越标准指标和大模型评估表现
  • 特别适合大规模学生模型的强化学习知识蒸馏场景

我们提出一种新型的多步回报估计方法(称为KETCHUP),用于基于强化学习的知识蒸馏(KD)在文本生成任务中的应用。核心思想是通过贝尔曼最优方程构造K步回报。理论分析表明,该多步形式可降低梯度估计方差,从而在学生模型规模较大时显著改善强化学习优化效果。在三个文本生成任务上的实证评估显示,该方法在标准任务指标和基于大语言模型(LLM)的评估中均取得更优性能。结果表明,多步回报诱导为提升大语言模型研究中的强化学习知识蒸馏提供了有前景的新方向。

原文摘要 · Abstract (English)

We propose a novel k-step return estimation method (called KETCHUP) for Reinforcement Learning(RL)-based knowledge distillation (KD) in text generation tasks. Our idea is to induce a K-step return by using the Bellman Optimality Equation for multiple steps. Theoretical analysis shows that this K-step formulation reduces the variance of the gradient estimates, thus leading to improved RL optimization especially when the student model size is large. Empirical evaluation on three text generation tasks demonstrates that our approach yields superior performance in both standard task metrics and large language model (LLM)-based evaluation. These results suggest that our K-step return induction offers a promising direction for enhancing RL-based KD in LLM research.

强化学习知识蒸馏大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。