arXiv:2502.01100cs.AIcs.CL2025-02ICML被引 208

测试大模型在复杂逻辑题中的表现,发现越难题准确率越低。

ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning

  • 用可控复杂度的逻辑谜题框架评估模型推理能力
  • 复杂度上升时准确率显著下降,大模型也难突破
  • 提出采样、回溯等策略提升逻辑推理效果

我们研究了大语言模型(LLMs)在复杂非单调逻辑推理中的能力及其可扩展性。为此,提出ZebraLogic,一个基于约束满足问题(CSPs)生成逻辑网格谜题的综合性评估框架。该框架支持生成具有可控且可量化的复杂度的谜题,系统研究Llama、o1模型和DeepSeek-R1等模型在难度递增下的推理表现。通过涵盖广泛的搜索空间复杂度和多样逻辑约束,ZebraLogic构建了结构化环境以评估推理能力。结果表明,随着问题复杂度增加,准确率出现显著下降——我们称之为‘复杂度诅咒’。这一局限即使在更大模型和更高推理计算量下依然存在,表明当前LLM推理能力存在固有瓶颈。此外,我们探索了多种增强策略,包括Best-of-N采样、回溯机制和自验证提示。研究揭示了LLM推理的可扩展性边界,指出了根本性限制,并为改进方向提供依据。

原文摘要 · Abstract (English)

We investigate the logical reasoning capabilities of large language models (LLMs) and their scalability in complex non-monotonic reasoning. To this end, we introduce ZebraLogic, a comprehensive evaluation framework for assessing LLM reasoning performance on logic grid puzzles derived from constraint satisfaction problems (CSPs). ZebraLogic enables the generation of puzzles with controllable and quantifiable complexity, facilitating a systematic study of the scaling limits of models such as Llama, o1 models, and DeepSeek-R1. By encompassing a broad range of search space complexities and diverse logical constraints, ZebraLogic provides a structured environment to evaluate reasoning under increasing difficulty. Our results reveal a significant decline in accuracy as problem complexity grows -- a phenomenon we term the curse of complexity. This limitation persists even with larger models and increased inference-time computation, suggesting inherent constraints in current LLM reasoning capabilities. Additionally, we explore strategies to enhance logical reasoning, including Best-of-N sampling, backtracking mechanisms, and self-verification prompts. Our findings offer critical insights into the scalability of LLM reasoning, highlight fundamental limitations, and outline potential directions for improvement.

逻辑推理大模型评测可扩展性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。