arXiv:2509.01236cs.CLcs.AI2025-09

揭示思维链推理中上下文学习与预训练先验的相互作用机制

Rethinking the Chain-of-Thought: The Roles of In-Context Learning and Pre-trained Priors

  • 从词汇级分析推理过程,发现模型依赖预训练先验
  • 充足示例能引导模型转向上下文信号,错误提示则导致不稳定
  • 长思维链提示可激发更长推理链,提升下游任务表现

思维链推理已成为提升模型推理能力的关键方法。尽管关注度日益提高,其内在机制仍不明确。本文从上下文学习与预训练先验的双重关系出发,深入探究思维链推理的工作机制。首先在词汇层面细粒度分析推理路径,考察模型推理行为;其次通过逐步引入噪声示例,研究模型如何在预训练先验与错误上下文信息间权衡;最后检验提示工程能否诱导大模型进行深度思考。大量实验揭示三个关键发现:(1) 模型不仅快速掌握词汇层面的推理结构,还能理解深层逻辑模式,但严重依赖预训练先验;(2) 提供足够示例可使模型决策从预训练先验转向上下文信号,而误导性提示则引发不稳定性;(3) 长思维链提示能诱导模型生成更长的推理链,从而提升下游任务性能。

原文摘要 · Abstract (English)

Chain-of-Thought reasoning has emerged as a pivotal methodology for enhancing model inference capabilities. Despite growing interest in Chain-of-Thought reasoning, its underlying mechanisms remain unclear. This paper explores the working mechanisms of Chain-of-Thought reasoning from the perspective of the dual relationship between in-context learning and pretrained priors. We first conduct a fine-grained lexical-level analysis of rationales to examine the model's reasoning behavior. Then, by incrementally introducing noisy exemplars, we examine how the model balances pretrained priors against erroneous in-context information. Finally, we investigate whether prompt engineering can induce slow thinking in large language models. Our extensive experiments reveal three key findings: (1) The model not only quickly learns the reasoning structure at the lexical level but also grasps deeper logical reasoning patterns, yet it heavily relies on pretrained priors. (2) Providing sufficient exemplars shifts the model's decision-making from pretrained priors to in-context signals, while misleading prompts introduce instability. (3) Long Chain-of-Thought prompting can induce the model to generate longer reasoning chains, thereby improving its performance on downstream tasks.

思维链大模型推理提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。