提出新方法区分大模型推理与记忆,发现主流模型实际依赖记忆而非真推理。
None of the Others: a General Technique to Distinguish Reasoning from Memorization in Multiple-Choice LLM Evaluation Benchmarks
- 设计全新题型变异技术,让正确答案不再依赖已见词或概念。
- 所有模型在新测试中准确率平均下降57%(MMLU)和50%(UNED-Access 2024)。
- 揭示当前评测存在数据污染,适合关注模型真实推理能力的研究者阅读。
在大模型评估中,通常通过数值变化来区分推理与记忆。本文提出一种适用于多选题的通用变异方法,使正确答案与过往见过的词汇或概念完全脱钩,迫使模型必须真正理解与推理才能作答。我们使用该方法在英文和西班牙语的两个数据集上评估了最先进的专有与开源模型:公开的MMLU基准和私有的UNED-Access 2024数据集。结果显示,所有模型在新测试中准确率均显著下降,平均损失分别为57%(MMLU)和50%(UNED-Access 2024),不同模型降幅在10%至93%之间。值得注意的是,表现最佳的模型(OpenAI-o3-mini)并非最稳健的(DeepSeek-R1-70B),说明标准评测中表现优异的模型未必具备更强的推理能力。此外,公开数据集和原始语言题目准确率下降更明显,表明当前模型存在数据污染问题,并凸显记忆在回答中的重要作用。
原文摘要 · Abstract (English)
In LLM evaluations, reasoning is often distinguished from recall/memorization by performing numerical variations to math-oriented questions. Here we introduce a general variation method for multiple-choice questions that completely dissociates the correct answer from previously seen tokens or concepts, requiring LLMs to understand and reason (rather than memorizing) in order to answer correctly. Using this method, we evaluate state-of-the-art proprietary and open-source LLMs on two datasets available in English and Spanish: the public MMLU benchmark and the private UNED-Access 2024 dataset. Results show that all models experience remarkable accuracy drops under our proposed variation, with an average loss of 57% on MMLU and 50% on UNED-Access 2024, ranging from 10% to 93% across models. Notably, the most accurate model in our experimentation (OpenAI-o3-mini) is not the most robust (DeepSeek-R1-70B), suggesting that the best models in standard evaluations may not be the ones with better reasoning capabilities. Also, we see larger accuracy drops in public (vs private) datasets and questions posed in their original language (vs a manual translation), which are signs of contamination and also point to a relevant role of recall/memorization in current LLMs' answers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。