arXiv:2511.05184cs.CL2025-11被引 6

用思维链提升小模型推理能力,效果显著

Effectiveness of Chain-of-Thought in Distilling Reasoning Capability from Large Language Models

  • 用思维链数据在白盒知识蒸馏中传递大模型推理能力
  • 蒸馏后的小模型在BBH硬基准上平均性能提升明显
  • 适合想高效压缩大模型推理能力的研究者

思维链(Chain-of-Thought, CoT)提示是提升大语言模型推理能力的常用方法。近期,CoT被用于知识蒸馏(KD),将大模型的推理能力迁移到小模型。本文通过白盒知识蒸馏实验,研究CoT在从Qwen和Llama2系列大模型向小模型迁移推理能力中的作用。实验使用CoT-Collection数据集生成的CoT数据,蒸馏后的模型在BIG-Bench-Hard(BBH)基准的自然语言推理与理解任务上进行评估。结果表明,引入CoT能有效提升白盒知识蒸馏的效果,使小模型在BBH任务上的平均表现显著改善。

原文摘要 · Abstract (English)

Chain-of-Thought (CoT) prompting is a widely used method to improve the reasoning capability of Large Language Models (LLMs). More recently, CoT has been leveraged in Knowledge Distillation (KD) to transfer reasoning capability from a larger LLM to a smaller one. This paper examines the role of CoT in distilling the reasoning capability from larger LLMs to smaller LLMs using white-box KD, analysing its effectiveness in improving the performance of the distilled models for various natural language reasoning and understanding tasks. We conduct white-box KD experiments using LLMs from the Qwen and Llama2 families, employing CoT data from the CoT-Collection dataset. The distilled models are then evaluated on natural language reasoning and understanding tasks from the BIG-Bench-Hard (BBH) benchmark, which presents complex challenges for smaller LLMs. Experimental results demonstrate the role of CoT in improving white-box KD effectiveness, enabling the distilled models to achieve better average performance in natural language reasoning and understanding tasks from BBH.

思维链知识蒸馏推理能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。