arXiv:2504.02789cs.CL2025-04Conference of the …被引 8

LLM记忆力超人但控制力弱,难靠记忆解决复杂问题。

Strong Memory, Weak Control: An Empirical Study of Executive Functioning in LLMs

  • 用经典记忆任务测试LLM,发现其记忆容量普遍高于人类。
  • 记忆能力强但无法提升推理或执行任务表现,说明控制力不足。
  • 适合关注模型认知缺陷与人类智能差异的研究者阅读。

工作记忆(working memory)是人类智力与执行功能的关键组成部分,与流体智力(包括推理与问题解决)密切相关。我们采用一系列经典工作记忆任务来评估大语言模型(LLMs)的工作记忆能力。结果显示,在多数情况下,LLMs的表现超过人类正常水平。然而,工作记忆容量的提升并未带来其他执行功能任务或问题解决基准上的性能改善。这表明LLMs可能存在注意力控制和认知灵活性缺陷,难以抑制自动反应或适应信息变化。当前的推理模型在弥补这些缺陷方面效果有限。

原文摘要 · Abstract (English)

Working memory, or the ability to hold and manipulate information in the mind, is a critical component of human intelligence and executive functioning. It is correlated with performance on various cognitive tasks, including measures of fluid intelligence, which encompasses reasoning and problem solving. We use a comprehensive set of classic working memory tasks to estimate the working memory capacity of large language models (LLMs). We find that in most cases, LLMs exceed normative human scores. However, we do not find that the increased capacity of working memory is associated with higher performance on other executive functioning tasks or problem solving benchmarks. These results suggest that LLMs may have deficits in attentional control and cognitive flexibility, which result in difficulties with inhibiting automatic responses and adapting to shifting information. Our findings suggest that current reasoning models have mixed results in compensating for these deficits.

大模型认知工作记忆执行功能智能缺陷

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。