评测大模型在LSAT逻辑题上的推理能力,发现其可通过反思改进错误。
Lost in the Logic: An Evaluation of Large Language Models' Reasoning Capabilities on LSAT Logic Games
- 构建了包含元数据的LSAT逻辑题数据集,用于评估大模型推理能力。
- GPT-4在反思提示下准确率达70%,GPT-3.5达46%,显著优于基础链式思维。
- 揭示模型在不同类型逻辑题中的表现差异及常见推理错误类型。
本文评估大型语言模型(LLMs)在法学院入学考试(LSAT)逻辑游戏部分的表现。该部分涉及复杂的逻辑推理任务,是检验现代大模型处理高难度逻辑推理能力的重要基准。作者构建了一个包含元数据的LSAT逻辑游戏数据集,并在链式思维提示(Chain-of-Thought prompting)设置下广泛评估模型表现。由于初始表现较弱,研究在小样本数据集上探索其他提示框架,借鉴反射机制(Reflexion)思想进行优化。结果显示,采用该方法后,GPT-4在该子集上的准确率达到70%,GPT-3.5为46%,显著提升,表明大模型具备修正自身逻辑错误的能力。最后,通过人类标注分析模型在不同题型上的表现优劣及典型错误类型,深入揭示了当前大模型的逻辑推理能力边界。
原文摘要 · Abstract (English)
In this thesis, I evaluate the performance of Large Language Models (LLMs) on the Law School Admissions Test (LSAT), specifically the Logic Games section of the test. I focus on this section because it presents a complex logical reasoning task and thus is a valuable source of data for evaluating how modern, increasingly capable LLMs can handle hard logical reasoning tasks. I construct a dataset of LSAT logic games and their associated metadata, and extensively evaluate LLMs' performance in a Chain-of-Thought prompting setting. Given the weak performance in this setting, I explore other prompting frameworks on a smaller subset of the dataset, adapting ideas from Reflexion to this task. This results in a substantially improved accuracy of 70 percent for GPT-4 and 46 percent for GPT-3.5 on this data subset, highlighting the capacity of LLMs to revise their logical errors, despite initially weak performance. Finally, I analyze the types of logic games that models perform better or worse on, as well as the types of logical errors I observe from human annotation, providing detailed insights on the logical reasoning capabilities of LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。