揭示大模型推理中错误累积机制,提出提升正确率的新思路
Rethinking External Slow-Thinking: From Snowball Errors to Probability of Correct Reasoning
- 用信息论分析推理错误雪球效应,解释为何逐步思考能减错
- 发现扩展搜索范围比复杂框架更有效,长期提升推理准确率
- 适合关注大模型推理可靠性与优化策略的研究者
测试时缩放(也称慢思考)已被证明可提升大语言模型在多步推理任务中的表现。然而,尽管广泛应用,其内在机制仍不清晰。本文从理论角度探究外部慢思考机制,首先分析大模型推理过程中的雪球错误效应,并通过信息论将其与正确推理概率关联。基于此,我们表明外部慢思考本质上是降低错误概率的策略。进一步对比了从简单到复杂的多种外部慢思考方法,揭示其差异与联系。研究发现,方法有效性并非由具体框架决定,扩大搜索范围或增强模型内部推理能力可能带来更持久的性能提升。代码已开源。
原文摘要 · Abstract (English)
Test-time scaling, which is also often referred to as slow-thinking, has been demonstrated to enhance multi-step reasoning in large language models (LLMs). However, despite its widespread utilization, the mechanisms underlying slow-thinking methods remain poorly understood. This paper explores the mechanisms of external slow-thinking from a theoretical standpoint. We begin by examining the snowball error effect within the LLM reasoning process and connect it to the likelihood of correct reasoning using information theory. Building on this, we show that external slow-thinking methods can be interpreted as strategies to mitigate the error probability. We further provide a comparative analysis of popular external slow-thinking approaches, ranging from simple to complex, highlighting their differences and interrelationships. Our findings suggest that the efficacy of these methods is not primarily determined by the specific framework employed, and that expanding the search scope or the model's internal reasoning capacity may yield more sustained improvements in the long term. We open-source our code at https://github.com/ZyGan1999/Snowball-Errors-and-Probability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。