arXiv:2505.15927stat.MLcs.LG2025-05NeurIPS被引 7

链式思维监督可显著降低学习样本需求,理论揭示其核心是推理信息增益。

CoT Information: Improved Sample Complexity under Chain-of-Thought Supervision

论文配图:CoT Information: Improved Sample Complexity under Chain-of-Thought Supervision
图 1 · 摘自论文原文
  • 引入链式思维信息度量,量化推理过程带来的额外判别能力
  • 证明链式思维下样本复杂度可降至标准方法的 d/I 倍,提升学习速度
  • 适用于需要多步推理的模型训练,为大模型可解释性提供理论支撑

标准输入输出监督在学习涉及多步推理的复杂函数时面临挑战。链式思维(CoT)监督通过提供中间推理步骤与最终输出,成为提升大语言模型推理能力的关键技术。本文建立CoT监督下的统计学习理论,核心在于区分训练目标(CoT风险)与测试目标(端到端风险)的差异。通过引入链式思维信息度量 $\mathcal{I}_{\mathcal{D}, h_\star}^{\mathrm{CoT}}(ε; \calH)$,该度量量化观察推理过程所获得的额外判别力,首次实现对两类风险的精确关联。理论表明,达到目标端到端误差 $ε$ 所需样本复杂度为 $d/\mathcal{I}_{\mathcal{D}, h_\star}^{\mathrm{CoT}}(ε; \calH)$,其中 $d$ 为假设类复杂度,远优于标准 $d/ε$ 率。同时建立了基于链式思维信息的信息论下界。结果表明,链式思维信息是此类学习任务的根本统计复杂度度量。

原文摘要 · Abstract (English)

Learning complex functions that involve multi-step reasoning poses a significant challenge for standard supervised learning from input-output examples. Chain-of-thought (CoT) supervision, which provides intermediate reasoning steps together with the final output, has emerged as a powerful empirical technique, underpinning much of the recent progress in the reasoning capabilities of large language models. This paper develops a statistical theory of learning under CoT supervision. A key characteristic of the CoT setting, in contrast to standard supervision, is the mismatch between the training objective (CoT risk) and the test objective (end-to-end risk). A central part of our analysis, distinguished from prior work, is explicitly linking those two types of risk to achieve sharper sample complexity bounds. This is achieved via the *CoT information measure* $\mathcal{I}_{\mathcal{D}, h_\star}^{\mathrm{CoT}}(ε; \calH)$, which quantifies the additional discriminative power gained from observing the reasoning process. The main theoretical results demonstrate how CoT supervision can yield significantly faster learning rates compared to standard E2E supervision. Specifically, it is shown that the sample complexity required to achieve a target E2E error $ε$ scales as $d/\mathcal{I}_{\mathcal{D}, h_\star}^{\mathrm{CoT}}(ε; \calH)$, where $d$ is a measure of hypothesis class complexity, which can be much faster than standard $d/ε$ rates. Information-theoretic lower bounds in terms of the CoT information are also obtained. Together, these results suggest that CoT information is a fundamental measure of statistical complexity for learning under chain-of-thought supervision.

链式思维学习理论样本效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。