arXiv:2504.00509cs.AIcs.CL2025-04中稿 · AACL-IJCNLP 2025被引 18

顶尖大模型在小学题上靠背答案,微调条件就大幅失分。

Recitation over Reasoning: How Cutting-Edge Language Models Can Fail on Elementary School-Level Reasoning Problems?

  • 设计新评测集检测模型是否死记硬背
  • 换一句条件,顶级模型准确率降60%
  • 提醒研究者警惕模型虚假智能

近年来,大语言模型评测任务从小学水平快速升级到前沿难题,让研究者误以为已接近超越人类智能。但这种表现究竟是真正的推理能力,还是仅凭训练中见过的海量网络内容机械复述?为此,我们提出RoR-Bench——一个新型多模态评测基准,用于检测模型在简单推理题中因条件细微变化而产生的复现行为。实证分析显示,当前顶尖大模型普遍存在严重复现现象:仅更改题干中一个短语,OpenAI-o1和DeepSeek-R1等模型在小学级算术与逻辑题上的性能便平均下降60%。这一发现警示学界,必须重新审视大模型的真实智能水平。

原文摘要 · Abstract (English)

The rapid escalation from elementary school-level to frontier problems of the difficulty for LLM benchmarks in recent years have weaved a miracle for researchers that we are only inches away from surpassing human intelligence. However, is the LLMs' remarkable reasoning ability indeed comes from true intelligence by human standards, or are they simply reciting solutions witnessed during training at an Internet level? To study this problem, we propose RoR-Bench, a novel, multi-modal benchmark for detecting LLM's recitation behavior when asked simple reasoning problems but with conditions subtly shifted, and conduct empirical analysis on our benchmark. Surprisingly, we found existing cutting-edge LLMs unanimously exhibits extremely severe recitation behavior; by changing one phrase in the condition, top models such as OpenAI-o1 and DeepSeek-R1 can suffer 60 percent performance loss on elementary school-level arithmetic and reasoning problems. Such findings are a wake-up call to the LLM community that compels us to re-evaluate the true intelligence level of cutting-edge LLMs.

大模型评估推理能力复现行为

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。