arXiv:2605.21260cs.LG2026-05

从学习理论角度揭示思维链的收益与代价,明确何时有用、何时有害。

On the Cost and Benefit of Chain of Thought: A Learning-Theoretic Perspective

  • 将思维链建模为答案映射与推理规则的交互,分解出收益与代价项
  • 代价由错误累积导致,即使答案接近真实,代价仍可能无限大
  • 在稳定条件下,误差增长呈线性或指数级,可精确控制

我们构建了一个学习理论框架来理解思维链(CoT)。将 CoT 模型化为答案映射与自回归生成中间问题的推理规则之间的交互,并定义了该交互下假设的推理风险。首个结果是该风险的紧致分解:一个捕获收益的预言轨迹风险(OTR),在领域自适应问题中退化为目标域风险;另一个捕获代价的轨迹不匹配风险(TMR),源于不匹配推理路径上的错误累积。我们证明,若损失函数、答案映射或推理规则任一缺乏稳定性,即便 OTR 为零且假设与真实值一致,TMR 仍可任意大。反之,在稳定性假设下,我们给出了 TMR 的紧致上界,其由一个精确放大因子决定,识别出有界、线性和指数误差增长三种模式。这些结果精确刻画了思维链何时有益、何时有害,以及二者之间的转换机制。

原文摘要 · Abstract (English)

We develop a learning-theoretic framework for understanding Chain of Thought (CoT). We model CoT as the interaction between an answer map and a chain rule that generates intermediate questions autoregressively, and define the reasoning risk of a hypothesis under this interaction. Our first result is a tight canonical decomposition of this risk into two terms with opposing roles: an oracle-trajectory risk (OTR), which captures the benefit of CoT and reduces to a target-domain risk in a domain adaptation problem, and a trajectory-mismatch risk (TMR), which captures the cost of CoT through error accumulation along mismatched reasoning trajectories. We then show that this cost is unavoidable without structure: if any one of the loss, the hypothesis answer map, or the chain rule lacks stability, the TMR can be arbitrarily large even when the OTR is zero and the hypothesis is uniformly close to the ground truth. Conversely, under stability, we prove a tight upper bound on the TMR governed by an exact amplification factor that identifies bounded, linear, and exponential error-growth regimes. Together, these results give a precise theory of when CoT helps, when it hurts, and what controls the transition between the two.

思维链学习理论误差分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。