arXiv:2603.16197cs.AIcs.CL2026-03

大模型在考试中表现好,可能只是记住了题库。

Are Large Language Models Truly Smarter Than Humans?

  • 用三种方法检测模型是否提前见过考题,发现13.8%题目被泄露
  • 模型在换种问法时准确率下降7-20个百分点,说明依赖记忆而非理解
  • 部分模型像‘背答案’一样答题,尤其在哲学等文科题上更明显

公开排行榜显示大语言模型在学术、法律和编程等领域已超越人类专家。然而,多数评测题库公开可查,存在模型训练数据与测试数据重合的系统性风险。本文对六款前沿模型(GPT-4o、GPT-4o-mini、DeepSeek-R1、DeepSeek-V3、Llama-3.3-70B、Qwen3-235B)开展三项互补实验:实验一采用词法污染检测管道分析MMLU 513道题,57个学科中整体污染率达13.8%(STEM类达18.1%,哲学最高达66.7%),对应准确率提升0.030至0.054;实验二对100道题进行改写与间接引用测试,平均准确率下降7.0个百分点,法律与伦理类下降达19.8个百分点;实验三使用TS-Guessing行为探针检测全部513题,72.5%触发远超随机水平的记忆信号,其中DeepSeek-R1呈现分布式记忆特征(76.6%部分重构,0%原样复现),解释其在实验二中的异常表现。三组实验一致得出污染程度排序:STEM > 专业类 > 社科 > 人文类。

原文摘要 · Abstract (English)

Public leaderboards increasingly suggest that large language models (LLMs) surpass human experts on benchmarks spanning academic knowledge, law, and programming. Yet most benchmarks are fully public, their questions widely mirrored across the internet, creating systematic risk that models were trained on the very data used to evaluate them. This paper presents three complementary experiments forming a rigorous multi-method contamination audit of six frontier LLMs: GPT-4o, GPT-4o-mini, DeepSeek-R1, DeepSeek-V3, Llama-3.3-70B, and Qwen3-235B. Experiment 1 applies a lexical contamination detection pipeline to 513 MMLU questions across all 57 subjects, finding an overall contamination rate of 13.8% (18.1% in STEM, up to 66.7% in Philosophy) and estimated performance gains of +0.030 to +0.054 accuracy points by category. Experiment 2 applies a paraphrase and indirect-reference diagnostic to 100 MMLU questions, finding accuracy drops by an average of 7.0 percentage points under indirect reference, rising to 19.8 pp in both Law and Ethics. Experiment 3 applies TS-Guessing behavioral probes to all 513 questions and all six models, finding that 72.5% trigger memorization signals far above chance, with DeepSeek-R1 displaying a distributed memorization signature (76.6% partial reconstruction, 0% verbatim recall) that explains its anomalous Experiment 2 profile. All three experiments converge on the same contamination ranking: STEM > Professional > Social Sciences > Humanities.

大模型评估记忆泄漏测评偏见

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。