arXiv:2502.07190cs.AI2025-02NAACL被引 5

分析大模型在抽象推理任务中的智能缺陷,揭示其难以应对新问题的本质原因。

Understanding LLMs' Fluid Intelligence Deficiency: An Analysis of the ARC Task

  • 通过控制实验解析大模型在ARC任务中的表现瓶颈
  • 发现三大局限:技能组合能力弱、不适应抽象输入、自左向右解码固有缺陷
  • 适合关注AI认知机制与通用智能评估的研究者

尽管大语言模型在众多自然语言任务中表现优异,但这些任务大多依赖于模型参数中编码的海量知识,而非在无先验知识的情况下解决新问题。在认知研究中,后者被称为流体智力,是衡量人类智能的关键能力。近期关于流体智力评估的研究已揭示大模型在此方面存在显著不足。本文通过控制实验,以最具代表性的ARC任务为例,分析大模型在展现流体智力时面临的挑战。研究发现现有大模型存在三大主要局限:技能组合能力有限、对抽象输入格式不熟悉,以及自左向右解码的内在缺陷。相关数据与代码可在 https://wujunjie1998.github.io/araoc-benchmark.github.io/ 获取。

原文摘要 · Abstract (English)

While LLMs have exhibited strong performance on various NLP tasks, it is noteworthy that most of these tasks rely on utilizing the vast amount of knowledge encoded in LLMs' parameters, rather than solving new problems without prior knowledge. In cognitive research, the latter ability is referred to as fluid intelligence, which is considered to be critical for assessing human intelligence. Recent research on fluid intelligence assessments has highlighted significant deficiencies in LLMs' abilities. In this paper, we analyze the challenges LLMs face in demonstrating fluid intelligence through controlled experiments, using the most representative ARC task as an example. Our study revealed three major limitations in existing LLMs: limited ability for skill composition, unfamiliarity with abstract input formats, and the intrinsic deficiency of left-to-right decoding. Our data and code can be found in https://wujunjie1998.github.io/araoc-benchmark.github.io/.

流体智力大模型认知缺陷ARC任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。