arXiv:2601.09113cs.AI2026-01被引 12

系统梳理大模型记忆机制,揭示其从静态到动态演进路径。

The AI Hippocampus: How Far are We From Human Memory?

  • 将记忆分为隐式、显式与代理三类框架,构建统一分类体系。
  • 指出显式记忆可动态更新知识,支持可查询的外部信息交互。
  • 适合关注AI长期记忆、多模态协同与自主智能的研究者阅读。

记忆在增强大语言模型(LLMs)和多模态大语言模型(MLLMs)的推理能力、适应性及上下文一致性方面起着基础作用。随着这些模型从静态预测器向具备持续学习与个性化推理能力的交互系统演进,记忆机制的引入成为其架构与功能发展的核心议题。本文综述了LLMs与MLLMs中的记忆研究,提出一个包含隐式、显式与代理记忆的分类体系。隐式记忆指预训练Transformer内部参数中嵌入的知识,涵盖记忆化、关联检索与上下文推理能力;近期研究致力于解析、操控与重构这种潜在记忆。显式记忆通过外部存储与检索组件扩展模型输出,支持文本语料、密集向量与图结构等动态可查询的知识表示,实现信息源的可扩展与可更新交互。代理记忆则在自主代理中引入持久化、时间延展的记忆结构,促进长期规划、自我一致性及多智能体协作行为,适用于具身与交互式AI场景。此外,综述还探讨了多模态环境中跨视觉、语言、音频与动作模态的连贯性记忆整合,讨论关键架构进展、基准任务及开放挑战,如记忆容量、对齐性、事实一致性与跨系统互操作性等问题。

原文摘要 · Abstract (English)

Memory plays a foundational role in augmenting the reasoning, adaptability, and contextual fidelity of modern Large Language Models and Multi-Modal LLMs. As these models transition from static predictors to interactive systems capable of continual learning and personalized inference, the incorporation of memory mechanisms has emerged as a central theme in their architectural and functional evolution. This survey presents a comprehensive and structured synthesis of memory in LLMs and MLLMs, organizing the literature into a cohesive taxonomy comprising implicit, explicit, and agentic memory paradigms. Specifically, the survey delineates three primary memory frameworks. Implicit memory refers to the knowledge embedded within the internal parameters of pre-trained transformers, encompassing their capacity for memorization, associative retrieval, and contextual reasoning. Recent work has explored methods to interpret, manipulate, and reconfigure this latent memory. Explicit memory involves external storage and retrieval components designed to augment model outputs with dynamic, queryable knowledge representations, such as textual corpora, dense vectors, and graph-based structures, thereby enabling scalable and updatable interaction with information sources. Agentic memory introduces persistent, temporally extended memory structures within autonomous agents, facilitating long-term planning, self-consistency, and collaborative behavior in multi-agent systems, with relevance to embodied and interactive AI. Extending beyond text, the survey examines the integration of memory within multi-modal settings, where coherence across vision, language, audio, and action modalities is essential. Key architectural advances, benchmark tasks, and open challenges are discussed, including issues related to memory capacity, alignment, factual consistency, and cross-system interoperability.

大模型记忆机制多模态智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。