arXiv:2505.23480cs.CL2025-05被引 23

通过自怀疑解析长思维链过思考问题,有效减少无效推理。

Revisiting Overthinking in Long Chain-of-Thought from the Perspective of Self-Doubt

  • 从自我怀疑角度量化分析过思考现象,发现冗余推理源于对答案的反复验证。
  • 新提示方法显著缩短回答长度,在多个数学任务上提升模型表现。
  • 适合关注推理效率与可信度的模型优化研究者参考。

推理大语言模型在复杂任务中表现优异,主要得益于长思维链(Long CoT)推理。然而,它们常出现过思考现象——即使已得出正确答案,仍进行不必要的推理步骤。以往工作多基于样本观察进行定性分析。本文从自我怀疑视角出发,定量分析过思考问题,发现过度重复验证正确答案是其主因。为此,提出一种简单有效的提示方法:先引导模型质疑输入问题的有效性,再基于评估结果简洁回应。在三个数学推理任务和四个含缺失前提的数据集上测试,该方法显著减少答案长度,并在四种主流推理模型上实现广泛性能提升。进一步分析表明,该方法有效减少推理步数,降低自我怀疑程度。

原文摘要 · Abstract (English)

Reasoning Large Language Models (RLLMs) have demonstrated impressive performance on complex tasks, largely due to the adoption of Long Chain-of-Thought (Long CoT) reasoning. However, they often exhibit overthinking -- performing unnecessary reasoning steps even after arriving at the correct answer. Prior work has largely focused on qualitative analyses of overthinking through sample-based observations of long CoTs. In contrast, we present a quantitative analysis of overthinking from the perspective of self-doubt, characterized by excessive token usage devoted to re-verifying already-correct answer. We find that self-doubt significantly contributes to overthinking. In response, we introduce a simple and effective prompting method to reduce the model's over-reliance on input questions, thereby avoiding self-doubt. Specifically, we first prompt the model to question the validity of the input question, and then respond concisely based on the outcome of that evaluation. Experiments on three mathematical reasoning tasks and four datasets with missing premises demonstrate that our method substantially reduces answer length and yields significant improvements across nearly all datasets upon 4 widely-used RLLMs. Further analysis demonstrates that our method effectively minimizes the number of reasoning steps and reduces self-doubt.

思维链模型推理自怀疑提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。