arXiv:2601.05019cs.CLcs.AI2026-01

大模型推理蒸馏会丢失人类认知成本,导致效果反而变差。

Hán Dān Xué Bù (Mimicry) or Qīng Chū Yú Lán (Mastery)? A Cognitive Perspective on Reasoning Distillation in Large Language Models

  • 用监督微调模仿教师推理路径,但失去认知结构。
  • 蒸馏后模型与人类难度相关性从0.64降至0.34,甚至不如原始版本。
  • 适合关注大模型本质推理机制的研究者阅读。

近期通过强化学习训练的大规模推理模型自然契合人类认知成本。然而我们发现,当前主流的推理蒸馏方法——通过监督微调(SFT)让学生模型模仿教师推理过程——无法传递这种认知结构。在14个模型上测试‘邯郸学步’(表面模仿)假设,结果表明蒸馏导致‘功能对齐坍缩’:教师模型与人类难度相关性为$ar{r}=0.64$,而蒸馏后的学生模型显著下降至$ar{r}=0.34$,且常低于自身预蒸馏基线(‘负迁移’)。分析表明,SFT引发‘活人崇拜’效应,学生仅形式化复制推理语言(冗长),未内化教师动态资源分配策略。因此,推理蒸馏使计算开销与认知需求脱钩,揭示人类级认知是主动强化的结果,而非被动模仿。

原文摘要 · Abstract (English)

Recent Large Reasoning Models trained via reinforcement learning exhibit a "natural" alignment with human cognitive costs. However, we show that the prevailing paradigm of reasoning distillation -- training student models to mimic these traces via Supervised Fine-Tuning (SFT) -- fails to transmit this cognitive structure. Testing the "Hán Dān Xué Bù" (Superficial Mimicry) hypothesis across 14 models, we find that distillation induces a "Functional Alignment Collapse": while teacher models mirror human difficulty scaling ($\bar{r}=0.64$), distilled students significantly degrade this alignment ($\bar{r}=0.34$), often underperforming their own pre-distillation baselines ("Negative Transfer"). Our analysis suggests that SFT induces a "Cargo Cult" effect, where students ritualistically replicate the linguistic form of reasoning (verbosity) without internalizing the teacher's dynamic resource allocation policy. Consequently, reasoning distillation decouples computational cost from cognitive demand, revealing that human-like cognition is an emergent property of active reinforcement, not passive imitation.

大模型推理认知模拟知识蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。