将神经科学记忆机制融入Transformer,提升模型长期记忆与持续学习能力
Memory-Augmented Transformers: A Systematic Review from Neuroscience Principles to Enhanced Model Architectures
- 从神经科学出发,构建记忆增强型Transformer的统一框架
- 揭示从静态缓存到测试时自适应学习的演进趋势
- 适合关注持续学习与认知启发模型的研究者
记忆是智能的核心,支撑生物与人工系统的学习、推理与适应能力。尽管Transformer在序列建模上表现优异,但在长程上下文保持、持续学习和知识整合方面存在局限。本文提出一个统一框架,连接神经科学中的动态多时标记忆、选择性注意与巩固机制,与内存增强Transformer的工程进展。我们从三大维度梳理最新成果:功能目标(上下文扩展、推理、知识整合、适应)、记忆表示(参数编码、状态驱动、显式存储、混合形式)和融合机制(注意力融合、门控控制、关联检索)。对核心记忆操作(读取、写入、遗忘、容量管理)的分析表明,系统正从静态缓存转向自适应、测试时学习的架构。仍存在可扩展性与干扰问题,新兴方案如分层缓冲与惊喜触发更新提供新解。本综述为迈向认知启发的终身学习型Transformer提供了路线图。
原文摘要 · Abstract (English)
Memory is fundamental to intelligence, enabling learning, reasoning, and adaptability across biological and artificial systems. While Transformer architectures excel at sequence modeling, they face critical limitations in long-range context retention, continual learning, and knowledge integration. This review presents a unified framework bridging neuroscience principles, including dynamic multi-timescale memory, selective attention, and consolidation, with engineering advances in Memory-Augmented Transformers. We organize recent progress through three taxonomic dimensions: functional objectives (context extension, reasoning, knowledge integration, adaptation), memory representations (parameter-encoded, state-based, explicit, hybrid), and integration mechanisms (attention fusion, gated control, associative retrieval). Our analysis of core memory operations (reading, writing, forgetting, and capacity management) reveals a shift from static caches toward adaptive, test-time learning systems. We identify persistent challenges in scalability and interference, alongside emerging solutions including hierarchical buffering and surprise-gated updates. This synthesis provides a roadmap toward cognitively-inspired, lifelong-learning Transformer architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。