提出新基准,测试大模型是否真能抽象推理而非死记硬背。
Benchmarking Abstract and Reasoning Abilities Through A Theoretical Perspective
- 用数学框架定义抽象推理:抓本质模式、用一致规则。
- 设计符号重映射实验,发现模型在非十进制运算上严重失效。
- 新指标\(\scoreDelta\)可有效检测模型是否依赖符号记忆而非抽象能力。
本文旨在建立一个简单、有效且理论严谨的基准,以严格评估大语言模型(LLMs)的抽象推理能力。为此,我们首先构建一个数学框架,将抽象推理定义为:(i) 提取独立于表面表征的本质模式;(ii) 对这些抽象模式应用一致规则。基于此框架,我们引入两个互补的新指标:\(\scoreGamma\) 衡量基础推理准确率,而 \(\scoreDelta\) 则量化模型对特定符号的依赖程度——这是区分真实抽象与单纯记忆的关键指标。为实施测量,我们设计了一个基准:在规则任务中系统性地进行符号重映射,迫使模型展现超越表面标记匹配的真实模式识别能力。对多种 LLM 的广泛评估(包括商用 API 模型、7B-70B 参数规模及多智能体系统)揭示:1)非十进制算术与符号推理存在显著缺陷;2)即使采用思维链提示,抽象差距依然存在;3)\(\scoreDelta\) 能稳健衡量记忆依赖性,尤其凸显了对操作数的特定记忆现象。结果表明,当前 LLM 尽管具备领域优势,但仍缺乏稳健的抽象推理能力,指明未来改进的关键方向。
原文摘要 · Abstract (English)
In this paper, we aim to establish a simple, effective, and theoretically grounded benchmark for rigorously probing abstract reasoning in Large Language Models (LLMs). To achieve this, we first develop a mathematic framework that defines abstract reasoning as the ability to: (i) extract essential patterns independent of surface representations, and (ii) apply consistent rules to these abstract patterns. Based on this framework, we introduce two novel complementary metrics: \(\scoreGamma\) measures basic reasoning accuracy, while \(\scoreDelta\) quantifies a model's reliance on specific symbols rather than underlying patterns - a key indicator of true abstraction versus mere memorization. To implement this measurement, we design a benchmark: systematic symbol remapping in rule-based tasks, which forces models to demonstrate genuine pattern recognition beyond superficial token matching. Extensive LLM evaluations using this benchmark (commercial API models, 7B-70B, multi-agent) reveal:1) critical limitations in non-decimal arithmetic and symbolic reasoning; 2) persistent abstraction gaps despite chain-of-thought prompting; and 3) \(\scoreDelta\)'s effectiveness in robustly measuring memory dependence by quantifying performance degradation under symbol remapping, particularly highlighting operand-specific memorization. These findings underscore that current LLMs, despite domain-specific strengths, still lack robust abstract reasoning, highlighting key areas for future improvement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。