arXiv:2505.20296cs.CLcs.AI2025-05被引 7

现有大模型推理像盲目漫游,缺乏系统探索能力。

Reasoning LLMs are Wandering Solution Explorers

  • 提出系统性解题标准,识别模型的无序探索缺陷。
  • 发现模型在复杂任务中错误率显著上升,且存在重复、幻觉等现象。
  • 建议评估推理过程结构,而非仅看最终答案。

大型语言模型(LLMs)通过测试时计算(TTC)技术如思维链提示和基于树的推理展现出强大推理能力。然而,本文认为当前推理型大模型(RLLMs)缺乏对解空间的系统性探索能力。论文形式化定义了系统性问题求解的标准,并揭示了常见失败模式,表明现有模型更像漫游者而非系统探索者。通过对多个前沿大模型的定性和定量分析,发现其普遍存在无效推理步骤、重复探索、幻觉或不忠实结论等问题。研究显示,尽管模型在简单任务上表现看似良好,但随着任务复杂度增加,性能急剧下降。基于此,论文倡导开发新指标与工具,不仅评估最终输出,更关注推理过程的结构本身。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated impressive reasoning abilities through test-time computation (TTC) techniques such as chain-of-thought prompting and tree-based reasoning. However, we argue that current reasoning LLMs (RLLMs) lack the ability to systematically explore the solution space. This paper formalizes what constitutes systematic problem solving and identifies common failure modes that reveal reasoning LLMs to be wanderers rather than systematic explorers. Through qualitative and quantitative analysis across multiple state-of-the-art LLMs, we uncover persistent issues: invalid reasoning steps, redundant explorations, hallucinated or unfaithful conclusions, and so on. Our findings suggest that current models' performance can appear to be competent on simple tasks yet degrade sharply as complexity increases. Based on the findings, we advocate for new metrics and tools that evaluate not just final outputs but the structure of the reasoning process itself.

大模型推理思维链评估方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。