Chain-of-Thought推理在上下文学习中反而表现更差,原因在于显式推理失效、隐式推理受干扰。
The Curse of CoT: On the Limitations of Chain-of-Thought in In-Context Learning
- 发现显式推理与隐式推理的混合机制导致CoT在模式型ICL中表现不佳
- 16个主流大模型在9个数据集上均显示CoT比直接回答效果差
- 即使长链条推理模型也难突破此局限,提示需重新设计推理方法
Chain-of-Thought(CoT)提示被广泛认为能增强大语言模型(LLMs)的推理能力。然而,我们的研究揭示了这一观点在基于模式的上下文学习(ICL)基础领域中的矛盾现象。通过对16个先进LLM和9个不同模式型ICL数据集的大量实验,我们发现CoT及其变体在不同模型规模和基准复杂度下始终劣于直接回答。为系统探究此反常现象,我们设计了多项实验验证多种假设。分析表明,驱动CoT在模式型ICL中表现的核心机制是显式-隐式推理的混合:显式推理因模型难以从示例中推断底层模式而失败,而隐式推理虽受CoT论证带来的上下文距离增加所干扰,仍可部分补偿并输出正确答案。这种混合机制解释了CoT相对性能下降的原因——弱显式推理产生的噪声破坏整体过程,即便隐式机制部分挽救结果。值得注意的是,尽管长链CoT模型在抽象与符号推理中表现优异,其高昂计算成本仍未完全克服这些限制。研究挑战了CoT普适有效的既有认知,为未来更精细、高效的推理方法提供了新方向。
原文摘要 · Abstract (English)
Chain-of-Thought (CoT) prompting has been widely recognized for its ability to enhance reasoning capabilities in large language models (LLMs). However, our study reveals a surprising contradiction to this prevailing perspective within the fundamental domain of pattern-based in-context learning (ICL). Through extensive experiments involving 16 state-of-the-art LLMs and nine diverse pattern-based ICL datasets, we demonstrate that CoT and its reasoning variants consistently underperform direct answering across varying model scales and benchmark complexities. To systematically investigate this unexpected phenomenon, we designed extensive experiments to validate several hypothetical explanations. Our analysis uncovers a fundamental hybrid mechanism of explicit-implicit reasoning driving CoT's performance in pattern-based ICL: while explicit reasoning falters due to LLMs' struggles to infer underlying patterns from demonstrations, implicit reasoning-disrupted by the increased contextual distance of CoT rationales-often compensates, delivering correct answers despite flawed rationales. This hybrid mechanism explains CoT's relative underperformance, as noise from weak explicit inference undermines the process, even as implicit mechanisms partially salvage outcomes. Notably, even long-CoT reasoning models, which excel in abstract and symbolic reasoning, fail to fully overcome these limitations despite higher computational costs. Our findings challenge existing assumptions regarding the universal efficacy of CoT, yielding novel insights into its limitations and guiding future research toward more nuanced and effective reasoning methodologies for LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。