arXiv:2505.19676cs.AI2025-05被引 2

前沿大模型推理能力在九个月内停滞,主要依赖提示工程而非真正进步。

Large Language Models' Reasoning Stalls: An Investigation into the Capabilities of Frontier Models

  • 通过分析模型对定理证明策略的使用,评估其逻辑推理能力。
  • 九个月间模型推理提升基本停滞,多数进展来自提示优化而非本质增强。
  • 模型最擅长自底向上的推理方式,但正确推理与正确答案关联弱。

本文研究了大语言模型(LLM)运用自动化定理证明(ATP)推理策略的能力。我们评估了2023年12月和2024年8月的前沿模型在PRONTOQA steamroller推理题上的表现。为此,我们开发了评估模型回答准确性和正确结论相关性的方法。结果表明,过去九个月内,模型推理能力的提升已基本停滞。通过追踪完成的标记数,我们发现自GPT-4发布以来的几乎所有进步,均可归因于隐藏系统提示或模型训练中对通用思维链提示策略的自动应用。在尝试的ATP推理策略中,当前前沿模型最擅长遵循自底向上(也称前向推理)策略。然而,模型回答中包含正确推理与最终得出正确结论之间仅存在微弱正相关。

原文摘要 · Abstract (English)

Empirical methods to examine the capability of Large Language Models (LLMs) to use Automated Theorem Prover (ATP) reasoning strategies are studied. We evaluate the performance of State of the Art models from December 2023 and August 2024 on PRONTOQA steamroller reasoning problems. For that, we develop methods for assessing LLM response accuracy and correct answer correlation. Our results show that progress in improving LLM reasoning abilities has stalled over the nine month period. By tracking completion tokens, we show that almost all improvement in reasoning ability since GPT-4 was released can be attributed to either hidden system prompts or the training of models to automatically use generic Chain of Thought prompting strategies. Among the ATP reasoning strategies tried, we found that current frontier LLMs are best able to follow the bottom-up (also known as forward-chaining) strategy. A low positive correlation was found between an LLM response containing correct reasoning and arriving at the correct conclusion.

大模型推理定理证明思维链

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。