强模型无需思维链示例,零样本推理反而更有效。
Revisiting Chain-of-Thought Prompting: Zero-shot Can Be Stronger than Few-shot
- 用零样本思维链替代传统示例,效果更好
- 增强版示例(大模型生成)仍无法提升性能
- 适合关注提示工程效率与模型本质能力的研究者
上下文学习(ICL)是大型语言模型的新兴能力,近期研究通过在ICL示例中引入思维链(CoT)来提升数学任务中的推理能力。然而,随着模型能力持续进步,现有强模型如Qwen2.5系列是否仍需传统CoT示例尚不明确。通过系统实验发现,对于这些强模型,添加传统CoT示例并未比零样本思维链(Zero-Shot CoT)带来性能提升,其主要作用仅是格式对齐。进一步测试以Qwen2.5-Max和DeepSeek-R1答案构建的增强型示例,结果表明其仍无法改善推理表现。深入分析显示,模型倾向于忽略示例内容,只关注指令,导致推理能力无实质提升。整体表明当前ICL+CoT框架在数学推理中存在局限,亟需重新审视上下文学习范式与示例定义。
原文摘要 · Abstract (English)
In-Context Learning (ICL) is an essential emergent ability of Large Language Models (LLMs), and recent studies introduce Chain-of-Thought (CoT) to exemplars of ICL to enhance the reasoning capability, especially in mathematics tasks. However, given the continuous advancement of model capabilities, it remains unclear whether CoT exemplars still benefit recent, stronger models in such tasks. Through systematic experiments, we find that for recent strong models such as the Qwen2.5 series, adding traditional CoT exemplars does not improve reasoning performance compared to Zero-Shot CoT. Instead, their primary function is to align the output format with human expectations. We further investigate the effectiveness of enhanced CoT exemplars, constructed using answers from advanced models such as \texttt{Qwen2.5-Max} and \texttt{DeepSeek-R1}. Experimental results indicate that these enhanced exemplars still fail to improve the model's reasoning performance. Further analysis reveals that models tend to ignore the exemplars and focus primarily on the instructions, leading to no observable gain in reasoning ability. Overall, our findings highlight the limitations of the current ICL+CoT framework in mathematical reasoning, calling for a re-examination of the ICL paradigm and the definition of exemplars.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。