揭示推理深度如何影响模型泛化能力的理论规律
An Asymptotic Theory of Chain-of-Thought in In-Context Learning

- 构建线性回归场景下的可解理论模型,分析推理深度与泛化误差关系
- 发现推理深度存在指数与多项式提升的临界点,过深会导致错误放大或饱和
- 验证了预训练数据和上下文信息丰富时长推理更有效,适合研究者参考
链式思维(CoT)推理已成为大语言模型生成多步推理过程的重要机制。然而,其泛化能力随推理深度的缩放行为仍不明确。本文针对线性回归中的上下文权重预测问题,建立一个可解析的理论模型,将测试时的推理表示为参数估计的迭代优化过程。基于高维渐近下的随机矩阵理论,推导出泛化误差随推理深度、预训练数据量和上下文长度变化的精确公式。分析揭示了从指数提升到多项式提升、饱和及过拟合之间的尖锐相变现象,并刻画了最优推理深度的依赖关系。结果表明,充分的预训练和上下文信息下,更深层次推理更有效;反之则易引发误差放大或饱和。实验在全学习线性注意力与Softmax注意力模型上验证了预测。研究提供了一个统一的理论框架,解释测试时推理深度对泛化的影响。
原文摘要 · Abstract (English)
Chain-of-thought (CoT) reasoning has become a widely used mechanism for eliciting multi-step reasoning in large language models by generating intermediate reasoning steps at inference time. Yet the scaling behavior of generalization with CoT depth remains poorly understood. To address this question, we study a theoretically solvable model of CoT for in-context weight prediction in linear regression, where test-time reasoning is represented as an iterative refinement of the weight-parameter estimate. Using tools from random matrix theory under high-dimensional asymptotics, we derive an exact formula for the generalization error as a function of reasoning depth, pretraining data amount, and context length. Our analysis reveals a sharp phase transition separating exponential and polynomial improvement, saturation, and overthinking, and characterizes how the optimal reasoning depth scales. We further show that deeper reasoning is most effective with sufficiently rich pretraining and in-context information, whereas limited pretraining or context makes longer reasoning prone to error amplification or saturation. We also validate these predictions through experiments on fully learned linear attention and softmax attention models. Our results provide a unified theoretical account of how test-time CoT depth affects generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。