arXiv:2503.16385cs.AI2025-03被引 40

提出DLCoT框架,优化长链推理的模型蒸馏效率

Deconstructing Long Chain-of-Thought: A Structured Reasoning Optimization Framework for Long CoT Distillation

  • 将长链推理分解为可处理片段并去除冗余解
  • 在非同源模型上提升蒸馏效果,性能显著改善
  • 适合追求高效推理能力的模型训练者

大语言模型在长链推理(Long CoT)方面展现出卓越推理能力。当前的R1蒸馏方案被视为一种低成本提升模型推理能力的有效方法,但其内在机制尚不明确。本研究考察了蒸馏数据的普适性,发现从Qwen-QwQ等教师模型中蒸馏长链推理能力,在非同源模型上效果显著下降,挑战了现有蒸馏方法的通用性假设。为深入理解长链推理的结构与模式,我们提出DLCoT(Deconstructing Long Chain-of-Thought)框架,包含三个关键步骤:(1) 数据分割,将复杂长链结构拆解;(2) 简化,剔除不可解和冗余解;(3) 中间错误状态优化。该方法显著提升模型性能与生成效率,助力高性能大模型开发。

原文摘要 · Abstract (English)

Recent advancements in large language models (LLMs) have demonstrated remarkable reasoning capabilities through long chain-of-thought (CoT) reasoning. The R1 distillation scheme has emerged as a promising approach for training cost-effective models with enhanced reasoning abilities. However, the underlying mechanisms driving its effectiveness remain unclear. This study examines the universality of distillation data and identifies key components that enable the efficient transfer of long-chain reasoning capabilities in LLM distillation. Our findings reveal that the effectiveness of long CoT reasoning distillation from teacher models like Qwen-QwQ degrades significantly on nonhomologous models, challenging the assumed universality of current distillation methods. To gain deeper insights into the structure and patterns of long CoT reasoning, we propose DLCoT (Deconstructing Long Chain-of-Thought), a distillation data enhancement framework. DLCoT consists of three key steps: (1) data segmentation to decompose complex long CoT structures, (2) simplification by eliminating unsolvable and redundant solutions, and (3) optimization of intermediate error states. Our approach significantly improves model performance and token efficiency, facilitating the development of high-performance LLMs.

链式推理模型蒸馏结构优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。