arXiv:2507.00711cs.LG2025-07中稿 · KONVENS 2025被引 8

大模型会过度思考,明明有正确答案还继续瞎推理。

Large Reasoning Models are not thinking straight: on the unreliability of thinking trajectories

  • 发现大模型在获得正确答案后仍会继续无效推理
  • 三款顶级模型在AIME2024上均出现纠错失败
  • 适合关注模型可解释性与可靠性的研究者

通过强化学习训练的大语言模型(LLMs)在推理基准测试中取得了显著进展。然而,越来越多的证据表明,这些模型常常生成更长但无效的思维链(CoTs),引发对基准成绩是否反映真实推理能力的质疑。本文揭示了‘过度思考’现象:即使明确给出正确解法,模型仍会无视并持续生成无意义的推理步骤,最终导致错误结论。在AIME2024数学基准测试中,对三款前沿模型的实验表明,它们整合修正信息的能力存在严重缺陷,为实现稳健且可解释的推理带来了新挑战。

原文摘要 · Abstract (English)

Large Language Models (LLMs) trained via Reinforcement Learning (RL) have recently achieved impressive results on reasoning benchmarks. Yet, growing evidence shows that these models often generate longer but ineffective chains of thought (CoTs), calling into question whether benchmark gains reflect real reasoning improvements. We present new evidence of overthinking, where models disregard correct solutions even when explicitly provided, instead continuing to generate unnecessary reasoning steps that often lead to incorrect conclusions. Experiments on three state-of-the-art models using the AIME2024 math benchmark reveal critical limitations in these models ability to integrate corrective information, posing new challenges for achieving robust and interpretable reasoning.

大模型推理可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。