arXiv:2412.10056cs.CLcs.AI2024-12被引 2

用高考真题测试大模型,发现高分不等于真能力强。

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs?

  • 用中国高考真题做闭卷测试,避免数据泄露影响。
  • 模型在不同难度题上表现异常稳定,且同难度题波动大。
  • 适合关注大模型真实能力评估的研究者和教育技术开发者。

大型语言模型(LLMs)通常通过人工设计的基准测试进行评估,普遍认为高分意味着更强的人类级表现。然而,有越来越多的担忧指出,模型可能因数据泄露而‘作弊’,在看似高分的同时却难以应对对人类而言简单的任务。为深入解决此问题,我们构建了基于中国高考(Gaokao)的综合性评估基准GAOKAO-Eval,并对发布于高考前的代表性模型进行‘闭卷’评估。结果表明,即使消除数据泄露并保证全面性,高分仍无法真正反映人类对齐的能力。为理解这一偏差,我们引入认知心理学中的Rasch模型分析模型得分模式,发现两个关键差异:1)模型在不同难度题目上表现异常一致;2)同难度题目间性能方差显著。此外,还发现教师对模型生成答案评分不一,且存在重复错误模式。这些现象与OpenAI o1的设计动机一致,而o1的‘推理即难度’机制可缓解该偏差。结果表明,GAOKAO-Eval能揭示现有基准未捕捉到的大模型能力局限,凸显了更契合大模型特性的难度分析的必要性。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are commonly evaluated using human-crafted benchmarks, under the premise that higher scores implicitly reflect stronger human-like performance. However, there is growing concern that LLMs may ``game" these benchmarks due to data leakage, achieving high scores while struggling with tasks simple for humans. To substantively address the problem, we create GAOKAO-Eval, a comprehensive benchmark based on China's National College Entrance Examination (Gaokao), and conduct ``closed-book" evaluations for representative models released prior to Gaokao. Contrary to prevailing consensus, even after addressing data leakage and comprehensiveness, GAOKAO-Eval reveals that high scores still fail to truly reflect human-aligned capabilities. To better understand this mismatch, We introduce the Rasch model from cognitive psychology to analyze LLM scoring patterns and identify two key discrepancies: 1) anomalous consistent performance across various question difficulties, and 2) high variance in performance on questions of similar difficulty. In addition, We identified inconsistent grading of LLM-generated answers among teachers and recurring mistake patterns. we find that the phenomenons are well-grounded in the motivations behind OpenAI o1, and o1's reasoning-as-difficulties can mitigate the mismatch. These results show that GAOKAO-Eval can reveal limitations in LLM capabilities not captured by current benchmarks and highlight the need for more LLM-aligned difficulty analysis.

大模型评估高考测试能力偏差认知建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。