构建复杂且无先验知识干扰的逻辑推理基准,精准评估大模型推理能力。
JustLogic: A Comprehensive Benchmark for Evaluating Deductive Reasoning in Large Language Models
- 合成生成高复杂度逻辑题,覆盖多样语言模式与论证结构。
- 顶尖模型表现接近人类平均但远低于人类极限,非推理模型更差。
- 可深入分析推理深度与论证形式对准确率的影响,适合模型评估者使用。
逻辑推理是大型语言模型(LLMs)的关键能力,近年来研究致力于提升其演绎推理水平。然而,现有演绎推理基准因任务复杂度不足、存在先验知识干扰以及错误分析肤浅而难以有效评估。为此,我们提出JustLogic——一个合成生成的演绎推理基准,用于严格评估LLMs。该基准具备三大特性:(i) 高度复杂,能生成多样化语言模式、词汇和论证结构;(ii) 无先验知识依赖,消除模型因已有知识带来的优势,确保仅通过演绎推理作答;(iii) 支持对推理深度与论证形式在模型准确率上的异质影响进行深入误差分析。实验结果表明:(i) 当前最先进(SOTA)推理类模型表现接近或优于人类平均水平,但显著低于人类上限;(ii) SOTA非推理模型仍低于人类平均水平。所有代码与数据可在 https://github.com/michaelchen-lab/JustLogic 获取。
原文摘要 · Abstract (English)
Logical reasoning is a critical component of Large Language Models (LLMs), and substantial research efforts in recent years have aimed to enhance their deductive reasoning capabilities. However, existing deductive reasoning benchmarks, which are crucial for evaluating and advancing LLMs, are inadequate due to their lack of task complexity, presence of prior knowledge as a confounder, and superficial error analysis. To address these deficiencies, we introduce JustLogic, a synthetically generated deductive reasoning benchmark designed for rigorous evaluation of LLMs. JustLogic is (i) highly complex, capable of generating a diverse range of linguistic patterns, vocabulary, and argument structures; (ii) prior knowledge independent, eliminating the advantage of models possessing prior knowledge and ensuring that only deductive reasoning is used to answer questions; and (iii) capable of in-depth error analysis on the heterogeneous effects of reasoning depth and argument form on model accuracy. Our experimental results on JustLogic reveal that (i) state-of-the-art (SOTA) reasoning LLMs perform on par or better than the human average but significantly worse than the human ceiling, and (ii) SOTA non-reasoning models still underperform the human average. All code and data are available at https://github.com/michaelchen-lab/JustLogic
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。