构建逻辑推理评测基准,揭示大模型在推理能力上的盲区。
Evaluating the Logical Reasoning Abilities of Large Reasoning Models
- 设计多类型逻辑推理评测集LogiEval,覆盖演绎、归纳等四类推理
- 模型在类比推理上超越人类,但整体能力不均衡且存在顽固短板
- 提出LogiEval-Hard子集,可有效诊断跨规模模型的逻辑瓶颈
大型推理模型通过在长链式思维数据上进行强化学习微调,在数学、编程及领域特定推理任务上达到顶尖水平。然而,其逻辑推理能力——人类认知的基础且独立于领域知识的能力——仍缺乏系统研究。为此,我们提出LogiEval,一个全面评估大模型逻辑推理能力的基准。该基准涵盖演绎、归纳、类比和溯因等多种推理类型,任务形式包括逻辑序列、论证分析等,数据源自高质量的人类考试(如LSAT、GMAT)。实验表明,现代模型在四选一论证分析题和类比推理中表现优异,甚至超过人类,但在不同推理类型和任务格式间表现不均,暴露出泛化能力的局限。分析发现,人类表现与模型失败模式不一致。为进一步推动研究,我们基于一种新筛选范式,构建了LogiEval-Hard——一个由小模型(Qwen3-30B-A3B)失败可靠预测的大模型难题集。结果显示,现代模型在LogiEval-Hard上存在显著且一致的失败现象。这表明逻辑推理的根本瓶颈在不同模型规模下持续存在,并确立了LogiEval-Hard作为诊断工具与严格测试平台的价值。
原文摘要 · Abstract (English)
Large reasoning models, often post-trained on long chain-of-thought (long CoT) data with reinforcement learning, achieve state-of-the-art performance on mathematical, coding, and domain-specific reasoning benchmarks. However, their logical reasoning capabilities - fundamental to human cognition and independent of domain knowledge - remain understudied. To address this gap, we introduce LogiEval, a holistic benchmark for evaluating logical reasoning in large reasoning models. LogiEval spans diverse reasoning types (deductive, inductive, analogical, and abductive) and task formats (e.g., logical sequence, argument analysis), sourced from high-quality human examinations (e.g., LSAT, GMAT). Our experiments demonstrate that modern reasoning models excel at 4-choice argument analysis problems and analogical reasoning, surpassing human performance, yet exhibit uneven capabilities across reasoning types and formats, highlighting limitations in their generalization. Our analysis reveals that human performance does not mirror model failure distributions. To foster further research, we curate LogiEval-Hard, a challenging subset identified through a novel screening paradigm where small-model failures (Qwen3-30B-A3B) reliably predict difficulties for larger models. Modern models show striking, consistent failures on LogiEval-Hard. This demonstrates that fundamental reasoning bottlenecks persist across model scales, and establishes LogiEval-Hard as both a diagnostic tool and a rigorous testbed for advancing logical reasoning in LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。