评测大模型在大学级推理中的真实理解能力,发现顶级模型仍靠猜测。
CLR-Bench: Evaluating Large Language Models in College-level Reasoning
- 构建涵盖16个学科的多类型题目集,含专家详解。
- 引入双指标评估:直接答题准确率63.31%,带理由回答仅39.00%。
- 揭示当前大模型缺乏真实推理能力,适合研究者与评测开发者参考。
大型语言模型(LLMs)在多项语言理解任务中表现出色。尽管已有新基准用于评估数学与计算机科学等领域的能力,但仅以多选题最终答案的准确率为衡量标准,难以验证模型对题目的真正理解。为填补这一空白,我们提出CLR-Bench,全面评估大模型在复杂大学级推理任务中的表现。具体而言:(i) 聚焦计算机科学与人工智能领域的16个挑战性大学学科,数据集包含5类题目,每道题配有专家级详细解析;(ii) 提出两个新指标进行公平评估:Q→A用于衡量直接答案预测性能,Q→AR则综合评估回答与提供推理理由的联合能力。我们在1,018道学科专属题目上对40个大模型进行了实验。结果表明,即使最先进的闭源模型GPT-4 turbo也倾向于‘猜测’答案:准确率从Q→A的63.31%骤降至Q→AR的39.00%,暴露出其推理能力严重不足。
原文摘要 · Abstract (English)
Large language models (LLMs) have demonstrated their remarkable performance across various language understanding tasks. While emerging benchmarks have been proposed to evaluate LLMs in various domains such as mathematics and computer science, they merely measure the accuracy in terms of the final prediction on multi-choice questions. However, it remains insufficient to verify the essential understanding of LLMs given a chosen choice. To fill this gap, we present CLR-Bench to comprehensively evaluate the LLMs in complex college-level reasoning. Specifically, (i) we prioritize 16 challenging college disciplines in computer science and artificial intelligence. The dataset contains 5 types of questions, while each question is associated with detailed explanations from experts. (ii) To quantify a fair evaluation of LLMs' reasoning ability, we formalize the criteria with two novel metrics. Q$\rightarrow$A is utilized to measure the performance of direct answer prediction, and Q$\rightarrow$AR effectively considers the joint ability to answer the question and provide rationale simultaneously. Extensive experiments are conducted with 40 LLMs over 1,018 discipline-specific questions. The results demonstrate the key insights that LLMs, even the best closed-source LLM, i.e., GPT-4 turbo, tend to `guess' the college-level answers. It shows a dramatic decrease in accuracy from 63.31% Q$\rightarrow$A to 39.00% Q$\rightarrow$AR, indicating an unsatisfactory reasoning ability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。