LLM记忆如薛定谔的猫,不查询就不确定是否存在。
Schrodinger's Memory: Large Language Models
- 用泛逼近定理解释大模型记忆机制
- 提出新方法评估不同模型的记忆能力
- 类比人脑记忆,揭示其量子态般不可观测性
记忆是人类所有活动的基础;没有记忆,人们几乎无法完成日常生活中的任何任务。随着大语言模型(LLMs)的发展,其语言能力日益接近人类水平。但这些模型是否具备记忆?基于现有表现,它们似乎表现出记忆特征。然而,关于其记忆机制的深层理论仍缺乏研究。本文利用泛逼近定理(UAT)解释了大模型中记忆的运作原理,并通过实验验证了多种大模型的记忆能力,提出一种基于记忆可检索性的新评估方法。我们主张,大模型的记忆具有类似薛定谔的猫的特性——仅在特定记忆被查询时才显现,否则处于不确定状态。此外,本文还对比了人脑与大模型的记忆机制,揭示其在运作逻辑上的异同。
原文摘要 · Abstract (English)
Memory is the foundation of all human activities; without memory, it would be nearly impossible for people to perform any task in daily life. With the development of Large Language Models (LLMs), their language capabilities are becoming increasingly comparable to those of humans. But do LLMs have memory? Based on current performance, LLMs do appear to exhibit memory. So, what is the underlying mechanism of this memory? Previous research has lacked a deep exploration of LLMs' memory capabilities and the underlying theory. In this paper, we use Universal Approximation Theorem (UAT) to explain the memory mechanism in LLMs. We also conduct experiments to verify the memory capabilities of various LLMs, proposing a new method to assess their abilities based on these memory ability. We argue that LLM memory operates like Schrödinger's memory, meaning that it only becomes observable when a specific memory is queried. We can only determine if the model retains a memory based on its output in response to the query; otherwise, it remains indeterminate. Finally, we expand on this concept by comparing the memory capabilities of the human brain and LLMs, highlighting the similarities and differences in their operational mechanisms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。