arXiv:2601.13992cs.CLcs.AI2026-01被引 1

用多教师协作提升小模型推理能力,避免错误模仿和知识遗忘。

"The Whole Is Greater Than the Sum of Its Parts": A Compatibility-Aware Multi-Teacher CoT Distillation Framework

  • 动态加权多个大模型的指导信号,依据学生实时适应度调整。
  • 在多个基准上超越现有方法,且不破坏原有知识结构。
  • 适合需要高效推理的小模型应用,如移动端部署。

链式思维(CoT)使大语言模型具备强大推理能力,但通常需巨大参数量。通过知识蒸馏将推理能力迁移至小型学生模型(SLMs)成为可行方案,但现有方法多依赖单一教师,受限于教师自身能力偏差与灾难性遗忘风险。尽管多教师方案更具吸引力,如何有效融合其指导信号仍具挑战:教师-学生不兼容可能放大幻觉,被动监督难以确保真实逻辑内化。为此,我们提出COMPACT框架,通过多维评估指标动态加权教师梯度:(1) 基于图的共识机制,识别主流推理路径以过滤误导性推理解释;(2) 基于互信息的适应性检测,捕捉学生真正理解推理过程的“顿悟时刻”;(3) 基于损失的难度评估,衡量学生对教师引导的接受程度,防止负向迁移。大量实验与隐空间分析表明,COMPACT有效整合多样推理能力,保持原模型知识结构完整,在多个基准上达到当前最优性能,同时显著缓解灾难性遗忘。

原文摘要 · Abstract (English)

Chain-of-Thought (CoT) reasoning empowers Large Language Models (LLMs) with remarkable capabilities but typically requires prohibitive parameter scales. CoT distillation has emerged as a promising paradigm to transfer reasoning prowess into compact Student Models (SLMs), but existing approaches often rely on a solitary teacher, capping the student's potential since individual LLMs often exhibit distinct capability biases and may suffer from catastrophic forgetting. While leveraging diverse teachers seems appealing, effectively fusing their supervisions remains challenging: teacher-student incompatibility risks amplifying hallucinations, and passive supervision fails to ensure genuine logic internalization. To address this, we introduce COMPACT, a framework that adaptively fuses supervisions from different teachers by dynamically weighting teacher gradients based on the student's real-time compatibility evaluated by a multi-dimensional metric: (1) Graph-based Consensus to filter misleading rationales by identifying mainstream reasoning paths; (2) Mutual-Information-based Adaptability to detect "epiphany moments" for genuinely understanding the reasoning process rather than merely imitating; and (3) Loss-based Difficulty to assess student receptivity to the teacher's guidance and prevent negative transfer. Extensive experiments and latent space analysis demonstrate that COMPACT effectively integrates diverse reasoning capabilities without damaging the model's original knowledge structure, achieving state-of-the-art performance on various benchmarks while mitigating catastrophic forgetting.

知识蒸馏链式思维多教师推理增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。