提出新评测框架,精准测试大模型在动态任务中的抽象推理能力
Truly Assessing Fluid Intelligence of Large Language Models through Dynamic Reasoning Evaluation
- 构建分层级动态推理评测集,含36个任务、4个认知水平
- 多数大模型在高阶认知上表现差,复杂度上升时泛化能力弱
- 适合关注模型真实推理能力评估的研究者和开发者
大型语言模型(LLMs)虽展现出类人推理能力,但其是否具备真正的流体智力(即在新情境中抽象推理与规则泛化的能力)仍存疑。现有评测基准或侧重领域知识(结晶智力),或缺乏可解释性。为此,我们提出DRE-Bench,一个基于分层认知框架的动态推理评测基准,包含36个抽象推理任务,按四个认知层次组织,每项任务设计多个动态变体以检验同一潜在规则。该设计实现对流体智力的细粒度、可解释、可靠评估。我们测试了多种前沿模型,包括通用型(GPT-4o、Claude 3.7)与专用推理模型(o1、DeepSeek-R1、QwQ、Skywork-OR1)。实验表明,尽管多数模型在低阶认知任务中表现良好且稳健,但在高阶认知任务中表现不佳,随着任务复杂度提升,泛化能力显著受限。研究揭示当前模型与真正人类流体智力之间的差距,并为系统追踪大模型推理进展提供新路径。
原文摘要 · Abstract (English)
Recent advances in large language models (LLMs) have demonstrated impressive reasoning capacities that mirror human-like thinking. However, whether LLMs possess genuine fluid intelligence (i.e., the ability to reason abstractly and generalize rules in novel situations) remains an open question. Existing reasoning benchmarks either focus on domain-specific knowledge (crystallized intelligence) or lack interpretability. To address these limitations, we propose DRE-Bench, a dynamic reasoning evaluation benchmark grounded in a hierarchical cognitive framework. DRE-Bench consists of 36 abstract reasoning tasks organized across four cognitive levels, with each task featuring multiple dynamic variants that test the same underlying latent rule. This design enables fine-grained, interpretable, and reliable assessments of fluid intelligence. We evaluate a range of state-of-the-art LLMs, including both general LLMs (GPT-4o, Claude 3.7) and reasoning LLMs (o1, DeepSeek-R1, QwQ, Skywork-OR1). Experimental results reveal that although most LLMs achieve competent and robust performance in low-level cognition, they struggle with high-level cognition and exhibit limited generalization as task complexity grows. Our findings highlight the gap between current LLMs and true human-like fluid intelligence and offer a new path for systematically tracking reasoning progress in LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。