arXiv:2510.13170cs.CL2025-10综述被引 4

用人类思维框架分析推理训练,让大模型更像人一样思考。

Putting on the Thinking Hats: A Survey on Chain of Thought Fine-tuning from the Perspective of Human Reasoning Mechanism

  • 以六顶思考帽为框架,系统分类推理训练方法。
  • 梳理主流数据集与模型表现,覆盖数学与代码任务。
  • 适合关注AI推理机制与认知科学交叉的研究者。

链式思维(Chain of Thought, CoT)微调旨在通过训练模型学习精心设计的推理轨迹,赋予大语言模型类人的推理能力,涵盖详细规划、发散思维、直觉判断、及时反思、内部思考和事实感知等。随着技术发展,大模型在数学推理和代码生成等任务上取得显著进步。然而,现有综述多聚焦技术细节,缺乏从人类认知机制出发的系统分析。为此,本文首次基于人类推理理论对CoT微调进行全面综述。受著名的六顶思考帽框架启发,我们将其作为分类工具,系统梳理并分析了多种微调方法。同时,基于该理论提出未来研究方向。此外,本文整理了现有数据集与模型性能对比,并维护一个实时更新的GitHub仓库(https://github.com/AI-Chen/Awesome-CoT-Finetuning),持续追踪领域进展。期望本综述能推动该快速发展的研究领域的创新与进步。

原文摘要 · Abstract (English)

Chain of thought (CoT) fine-tuning aims to endow large language models (LLMs) with reasoning capabilities by training them on curated reasoning traces. It leverages both supervised and reinforced fine-tuning to cultivate human-like reasoning skills in LLMs, including detailed planning, divergent thinking, intuitive judgment, timely reflection, internal thinking, and fact perception, etc. As CoT fine-tuning has advanced, LLMs have demonstrated substantial improvements in tasks such as mathematical reasoning and code generation. However, existing surveys about CoT fine-tuning primarily focus on technical aspects and overlook a systematic analysis from the perspective of human reasoning mechanisms. Given that the ultimate goal of CoT fine-tuning is to enable LLMs to reason like humans, it is crucial to investigate this technique through the lens of human cognition. To fill this gap, we present the first comprehensive survey of CoT fine-tuning grounded in human reasoning theory. Specifically, inspired by the well-known Six Thinking Hats framework, which systematically characterizes common human thinking modes using six metaphorical hats, we classify and examine CoT fine-tuning methods through this lens. Furthermore, building upon this theory, we outline potential directions for future research in CoT fine-tuning. In addition, we compile a comprehensive overview of existing datasets and model performances, and a real-time GitHub repository \footnote{https://github.com/AI-Chen/Awesome-CoT-Finetuning} that continuously tracks recent advances in this area is maintained. We hope this survey will serve as a valuable resource to inspire innovation and foster progress in this rapidly evolving field.

链式思维推理训练认知模型大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。