arXiv:2410.06547cs.CLcs.FL2024-10EMNLP被引 1

首个测试大模型计算理论推理能力的基准,涵盖4006道题

TuringQ: Benchmarking AI Comprehension in Theory of Computation

  • 构建4006道理论计算机科学题目,分四级难度覆盖7个核心领域
  • GPT-4在链式思考提示下表现最优,人类评估与自动评分结果高度一致
  • 微调Llama3-8B可提升计算推理能力,并迁移至代数等新任务

我们提出TuringQ,首个用于评估大语言模型(LLMs)在计算理论推理能力方面的基准。TuringQ包含4,006对本科与研究生水平的问答,按难度分为四个等级,覆盖七个核心理论领域。我们采用链式思考提示和专家人类评估,对多个开源LLM及GPT-4进行评测。此外,我们提出一种基于LLM的自动化评估系统,在准确性上与人工评估相当。将Llama3-8B在TuringQ上微调后,其推理能力显著提升,并在代数等跨领域任务中表现出增强性能。TuringQ既作为基准,也作为提升复杂计算推理能力的资源。我们的分析揭示了大模型在理论计算机科学理解上的能力边界与进展。

原文摘要 · Abstract (English)

We present TuringQ, the first benchmark designed to evaluate the reasoning capabilities of large language models (LLMs) in the theory of computation. TuringQ consists of 4,006 undergraduate and graduate-level question-answer pairs, categorized into four difficulty levels and covering seven core theoretical areas. We evaluate several open-source LLMs, as well as GPT-4, using Chain of Thought prompting and expert human assessment. Additionally, we propose an automated LLM-based evaluation system that demonstrates competitive accuracy when compared to human evaluation. Fine-tuning a Llama3-8B model on TuringQ shows measurable improvements in reasoning ability and out-of-domain tasks such as algebra. TuringQ serves as both a benchmark and a resource for enhancing LLM performance in complex computational reasoning tasks. Our analysis offers insights into LLM capabilities and advances in AI comprehension of theoretical computer science.

计算理论大模型评测推理能力链式思考

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。