解释了推理长度为何收敛,揭示了模型过拟合与欠拟合的权衡机制。
Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought
- 将推理过程视为连续语义空间中的优化问题,而非离散分步预测。
- 发现推理长度收敛是过拟合与欠拟合权衡的自然结果。
- 为基于强化学习的测试时扩展提供理论支持,适合研究推理机制者参考。
测试时扩展,尤其是通过强化学习实现的多步链式思维(CoT)推理,已成为提升大语言模型推理能力的关键范式。然而,传统基于标记级别的分析无法捕捉推理层面的宏观动态。为此,我们提出 CoT-Space 理论框架,将推理过程从离散标记预测重构为连续推理语义空间中的优化过程。通过从噪声与风险角度建模推理轨迹,并重用经典学习理论的基本原理,我们证明推理长度趋于最优是过拟合与欠拟合权衡的自然结果。实验中进一步利用强化学习验证上述结论。研究为基于强化学习的内部测试时扩展提供了机制性解释,为现代大模型的推理轨迹优化奠定了原则性理论基础。
原文摘要 · Abstract (English)
Test-time scaling, primarily manifested through multi-step Chain-of-Thought (CoT) reasoning via Reinforcement Learning (RL), has emerged as a pivotal paradigm for enhancing the reasoning capabilities of Large Language Models (LLMs). However, a significant theoretical gap persists: traditional token-level analysis fails to capture the macroscopic dynamics of reasoning-level scaling. To address this, we introduce CoT-Space, a novel theoretical framework that recasts the reasoning process from a discrete token-prediction task to an optimization process within a continuous, reasoning-level semantic space. By modeling the reasoning trajectory from both noise and risk perspectives and revitalizing foundational principles from classical learning theory, we demonstrate that the observed convergence to an optimal CoT length is a natural consequence of the fundamental trade-off between underfitting and overfitting. We further utilize RL as a tool to elicit and verify these results in our experiments. Our findings provide a mechanistic explanation for the internal test-time scaling via RL, offering a principled theoretical foundation to optimize reasoning trajectories in modern LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。