发现大模型靠死记硬背提升成绩,提出新评测方法分离真实能力与记忆。
Large Language Models Could Be Rote Learners
- 将选择题改造为三元知识结构,减少记忆干扰。
- 实验显示主流模型平均19.6%的知识依赖死记硬背。
- 适合关注评测可靠性与模型真实能力的研究者。
基于基准的评估(如多项选择题和开放问答)广泛用于大语言模型(LLMs)评价,但其可靠性受基准污染影响。当模型在训练中提前接触测试基准时,能力较弱的LLM会表现出虚高性能,导致评估结果失真。本研究将污染视为学习的固有部分,旨在分离真实能力获取与表面记忆。通过分析不同记忆条件下多项选择题的表现,我们发现一个反直觉现象:模型在记忆过的基准上表现反而更差,表明存在死记硬背与真实能力学习并存的现象。为此,我们提出TrinEval——一种新的评估框架,将多项选择题重构为以知识为中心的三元形式,降低记忆影响同时保留内在知识,从而在存在记忆的情况下评估真实能力。大量实验验证了TrinEval在重构基准上的有效性和鲁棒性,评估结果进一步揭示,主流大模型在MMLU和GSM8K数据集上平均有19.6%的知识点依赖死记硬背。
原文摘要 · Abstract (English)
Benchmark-based evaluation, e.g., multiple-choice questions (MCQs) and open-ended questions (OEQs), is widely used for evaluating Large Language Models (LLMs), yet their reliability is undermined by benchmark contamination. When pre-exposed to the testing benchmark during training, less capable LLMs have been found to achieve inflated performance, thereby yielding erroneous results in LLM evaluation. In this study, we reframe contamination as an inherent aspect of learning and seek to disentangle and expose genuine capability acquisition from superficial memorization in LLM evaluation. Following this, firstly, by analyzing model performance under different memorization conditions of MCQs, we uncover a counterintuitive trend: LLMs perform worse on memorized benchmarks than on non-memorized ones, indicating the coexistence of two learning phenomena, i.e., rote memorization and genuine capability learning. To disentangle them, we propose TrinEval, a novel evaluation framework that reformulates MCQs into an alternative knowledge-centric trinity format, reducing memorization while preserving inherent knowledge, enabling the evaluation of genuine capability in the presence of memorization. Extensive experiments validate the effectiveness and robustness of TrinEval in reformulating benchmarks, and the evaluation results further reveal that mainstream LLMs rely on rote memorization for an average of 19.6% of knowledge points across the MMLU and the GSM8K dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。