用可调控复杂度的逻辑题测试大模型推理能力
On the logical skills of large language models: evaluations using arbitrarily complex first-order logic problems
- 构建可控制复杂度的一阶逻辑题生成方法
- 多模型在集合论逻辑题上表现均不佳,难度越高越差
- 适合评估大模型逻辑推理能力,尤其对算法研究者
我们提出一种生成一阶逻辑语句的方法,其复杂度可在多个维度上精确控制。基于此方法,我们自动生成多个数据集,包含要求判断一阶逻辑语句在策梅洛-弗兰克尔集合论中真假的问题。这些问题的解答仅需掌握一阶逻辑和集合论的基本符号,但需要规划与逻辑推理能力,且可通过生成语句的复杂度任意提升难度。我们对多种大语言模型(包括 DeepSeek-R1 和 OpenAI o3-mini)进行了广泛评估。所有数据集、生成代码及评估结果均已公开于 https://github.com/bkuckuck/logical-skills-of-llms。
原文摘要 · Abstract (English)
We present a method of generating first-order logic statements whose complexity can be controlled along multiple dimensions. We use this method to automatically create several datasets consisting of questions asking for the truth or falsity of first-order logic statements in Zermelo-Fraenkel set theory. While the resolution of these questions does not require any knowledge beyond basic notation of first-order logic and set theory, it does require a degree of planning and logical reasoning, which can be controlled up to arbitrarily high difficulty by the complexity of the generated statements. Furthermore, we do extensive evaluations of the performance of various large language models, including recent models such as DeepSeek-R1 and OpenAI's o3-mini, on these datasets. All of the datasets along with the code used for generating them, as well as all data from the evaluations is publicly available at https://github.com/bkuckuck/logical-skills-of-llms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。