arXiv:2511.21743cs.CL2025-11

模型学任务时,先靠推理逐步变强,最后不用推理也能解题。

Scaling Competence, Shrinking Reasoning: Cognitive Signatures in Language Model Learning

  • 用人类学习四阶段理论分析模型推理过程
  • 推理长度先增后减,峰值在掌握任务时
  • 移除推理后性能不变,说明推理是临时支架

我们分析了语言模型在特定任务微调中的推理行为,并将其与人类工作记忆进行类比。基于认知科学,我们将训练动态对应到学习的四个阶段:模型初期无推理输出错误,随后开始推理但仍失败,进而有效推理,最终无需显式推理即可解决问题。我们发现,随着性能提升,推理标记长度增加,在‘有意识的熟练’阶段达到峰值,之后随模型内化任务而下降。值得注意的是,训练完成后,即使移除推理,模型仍保持性能——表明推理仅起辅助学习作用,而非持续必要。该演进过程提供了实用洞见:推理标记动态可作为诊断训练阶段、识别收敛和指导早停的信号。我们提出了追踪这一轨迹的指标,主张推理行为对理解与优化推理型模型训练具有重要价值。

原文摘要 · Abstract (English)

We analyze reasoning in language models during task-specific fine-tuning and draws parallel between reasoning tokens--intermediate steps generated while solving problem and the human working memory. Drawing from cognitive science, we align training dynamics with the Four Stages of Competence: models initially produce incorrect outputs without reasoning, then begin reasoning (but still fail), eventually reason effectively, and finally solve tasks without explicit reasoning. We find that reasoning token length expands as performance improves, peaks at the stage of conscious competence, then declines as the model internalizes the task. Notably, after training, models retain performance even when reasoning is removed--suggesting it scaffolded learning but is no longer needed. This progression offers actionable insights: reasoning token dynamics can serve as a signal for diagnosing training stage, identifying convergence, and guiding early stopping. We propose metrics to track this trajectory and argue that reasoning behavior is valuable for understanding and optimizing reasoning model training.

语言模型推理机制训练动态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。