arXiv:2409.12500cs.CLcs.AI2024-09中稿 · LERC COLING 2024被引 6

用大模型生成奖励信号,提升小模型知识蒸馏效果

LLMR: Knowledge Distillation with a Large Language Model-Induced Reward

  • 用大模型构建奖励函数指导小模型训练
  • 在对话生成与摘要任务中均超越传统蒸馏方法
  • 适合资源受限场景下部署高效语言模型

大型语言模型在自然语言处理任务中表现卓越,但计算成本高,难以在资源受限环境中部署。本文提出 LLMR,一种基于大语言模型诱导奖励函数的知识蒸馏方法。我们在对话生成和摘要任务的多个数据集上进行了实验。结果表明,LLMR 在不同任务和数据集上均持续优于传统知识蒸馏方法。

原文摘要 · Abstract (English)

Large language models have become increasingly popular and demonstrated remarkable performance in various natural language processing (NLP) tasks. However, these models are typically computationally expensive and difficult to be deployed in resource-constrained environments. In this paper, we propose LLMR, a novel knowledge distillation (KD) method based on a reward function induced from large language models. We conducted experiments on multiple datasets in the dialogue generation and summarization tasks. Empirical results demonstrate that our LLMR approach consistently outperforms traditional KD methods in different tasks and datasets.

知识蒸馏大模型语言模型强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。