研究大模型推理耗力如何随问题复杂度变化,发现存在临界点后耗力不再上升。
Reasoning Effort and Problem Complexity: A Scaling Analysis in LLMs
- 用可无限扩展的帐篷谜题测试模型推理耗力
- 问题复杂度超过临界点后,推理耗力不再增加甚至下降
- 揭示当前大模型推理能力的瓶颈,适合关注模型逻辑局限的研究者
大型语言模型(LLMs)展现出卓越的文本生成能力,近期训练范式的进展也带来了推理性能的突破。本文研究此类模型的推理努力如何随问题复杂度变化。我们采用具有已知线性时间解法的无限可扩展帐篷谜题(Tents puzzle),分析这一缩放行为。结果表明,推理努力随问题规模增长,但仅在达到临界复杂度前有效增加;超过该阈值后,推理努力不再上升,甚至可能下降。这一现象揭示了当前大模型在问题复杂度提升时逻辑连贯性的关键局限,强调了改进推理可扩展性的必要性。此外,我们的结果还揭示了当前先进推理模型在应对日益复杂的逻辑谜题时存在显著性能差异。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated remarkable text generation capabilities, and recent advances in training paradigms have led to breakthroughs in their reasoning performance. In this work, we investigate how the reasoning effort of such models scales with problem complexity. We use the infinitely scalable Tents puzzle, which has a known linear-time solution, to analyze this scaling behavior. Our results show that reasoning effort scales with problem size, but only up to a critical problem complexity. Beyond this threshold, the reasoning effort does not continue to increase, and may even decrease. This observation highlights a critical limitation in the logical coherence of current LLMs as problem complexity increases, and underscores the need for strategies to improve reasoning scalability. Furthermore, our results reveal significant performance differences between current state-of-the-art reasoning models when faced with increasingly complex logical puzzles.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。