大模型看似推理超强,实则复杂问题一上来就崩塌。
Reasoning Models Reason Well, Until They Don't
- 用可扩展复杂度的数据集测试大模型推理能力
- 模型在复杂度达标后性能骤降,无法泛化
- 现实世界中多数问题在模型成功区间,但长尾风险高
大型语言模型(LLMs)在推理任务上取得显著进展,但近期研究显示,当推理问题超出适度复杂度时,变压器和大模型会彻底失效。本文从大型推理模型(LRMs)的视角重新审视这一现象——即通过鼓励逐步推理和自验证进行微调的LLMs。尽管现有基准如NLGraph上表现惊人,甚至声称具备数学、物理、医学、法律等领域的通用推理与创新潜力,但通过更精细地增加问题复杂度,我们发现这些基准实际复杂度有限。为此,我们构建了新的深度推理数据集DeepRD,并设计生成流程以产生无限数量的可扩展复杂度样本。在此基础上评估模型在图连通性与自然语言证明规划中的表现,结果表明:当复杂度达到一定阈值时,LRM性能急剧下降且不具备泛化能力。我们还将该结果与真实世界知识图谱、交互图及证明数据集的复杂度分布进行对比,发现多数真实案例位于模型成功区间,但长尾部分暴露了巨大失败风险。分析揭示了LRMs的短期实用性,同时强调亟需能超越训练数据复杂度分布的新方法。
原文摘要 · Abstract (English)
Large language models (LLMs) have shown significant progress in reasoning tasks. However, recent studies show that transformers and LLMs fail catastrophically once reasoning problems exceed modest complexity. We revisit these findings through the lens of large reasoning models (LRMs) -- LLMs fine-tuned with incentives for step-by-step argumentation and self-verification. LRM performance on graph and reasoning benchmarks such as NLGraph seem extraordinary, with some even claiming they are capable of generalized reasoning and innovation in reasoning-intensive fields such as mathematics, physics, medicine, and law. However, by more carefully scaling the complexity of reasoning problems, we show existing benchmarks actually have limited complexity. We develop a new dataset, the Deep Reasoning Dataset (DeepRD), along with a generative process for producing unlimited examples of scalable complexity. We use this dataset to evaluate model performance on graph connectivity and natural language proof planning. We find that the performance of LRMs drop abruptly at sufficient complexity and do not generalize. We also relate our LRM results to the distributions of the complexities of large, real-world knowledge graphs, interaction graphs, and proof datasets. We find the majority of real-world examples fall inside the LRMs' success regime, yet the long tails expose substantial failure potential. Our analysis highlights the near-term utility of LRMs while underscoring the need for new methods that generalize beyond the complexity of examples in the training distribution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。