arXiv:2502.04667cs.LGcs.AI2025-02被引 6

CoT训练让大模型学会拆解问题,组合技能解决新难题。

Compositional Generalization from Learned Skills via CoT Training: A Theoretical and Structural Analysis for Reasoning

  • 通过分步推理,模型将简单技能组合成复杂能力
  • 在未知场景下仍能保持良好表现,收敛速度更快
  • 适合想提升模型推理鲁棒性的研究者和工程师

链式思维(CoT)训练显著提升了大语言模型的推理能力,但其增强泛化性的机制尚不清晰。本文揭示:组合泛化是核心——模型在CoT训练中系统性地组合已学简单技能以应对新且更复杂的任务。理论分析表明,信息论泛化界可分解为分布内(ID)与分布外(OOD)成分;非CoT模型因未见过组合模式而失败,而CoT模型通过技能组合实现强泛化。控制实验与真实场景验证显示,CoT训练加速收敛,提升从分布内到分布外的泛化能力,且对容许噪声保持稳健。结构分析发现,CoT训练将推理内化为两阶段组合电路,推理步骤数对应阶段数;相较于非CoT模型,其在浅层即完成中间结果求解,释放深层专注后续步骤。关键洞见:CoT训练教会模型如何思考——通过提供正确答案,培养组合推理能力,而非仅传递知识。本研究为设计高效CoT策略、增强模型推理鲁棒性提供重要参考。

原文摘要 · Abstract (English)

Chain-of-Thought (CoT) training has markedly advanced the reasoning capabilities of large language models (LLMs), yet the mechanisms by which CoT training enhances generalization remain inadequately understood. In this work, we demonstrate that compositional generalization is fundamental: models systematically combine simpler learned skills during CoT training to address novel and more complex problems. Through a theoretical and structural analysis, we formalize this process: 1) Theoretically, the information-theoretic generalization bounds through distributional divergence can be decomposed into in-distribution (ID) and out-of-distribution (OOD) components. Specifically, the non-CoT models fail on OOD tasks due to unseen compositional patterns, whereas CoT-trained models achieve strong generalization by composing previously learned skills. In addition, controlled experiments and real-world validation confirm that CoT training accelerates convergence and enhances generalization from ID to both ID and OOD scenarios while maintaining robust performance even with tolerable noise. 2) Structurally, CoT training internalizes reasoning into a two-stage compositional circuit, where the number of stages corresponds to the explicit reasoning steps during training. Notably, CoT-trained models resolve intermediate results at shallower layers compared to non-CoT counterparts, freeing up deeper layers to specialize in subsequent reasoning steps. A key insight is that CoT training teaches models how to think-by fostering compositional reasoning-rather than merely what to think, through the provision of correct answers alone. This paper offers valuable insights for designing CoT strategies to enhance LLMs' reasoning robustness.

链式思维推理泛化大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。