让大模型像人一样主动管理记忆,实现真正无限上下文。
Cognitive Workspace: Active Memory Management for LLMs -- An Empirical Study of Functional Infinite Context
- 模拟人类认知机制,主动筛选和管理记忆信息。
- 内存复用率58.6%,较传统方法提升显著(p<0.001)。
- 适合需要长期推理与复杂任务规划的场景。
大型语言模型(LLMs)虽已将上下文窗口扩展至数百万标记,但在上下文管理上仍面临根本性局限。本文提出认知工作区(Cognitive Workspace),一种超越传统检索增强生成(RAG)的新范式,借鉴了Baddeley的工作记忆模型、Clark的延伸心智假说与Hutchins的分布式认知框架。分析2024-2025年进展表明,尽管Infini-attention和StreamingLLM等技术实现了极长上下文,但缺乏元认知意识与主动规划能力。认知工作区通过三项创新解决:(1)主动记忆管理与信息精炼;(2)分层认知缓冲区支持持久工作状态;(3)任务驱动的上下文优化。实证结果显示,其平均内存复用率达58.6%(任务间54%-60%),较传统RAG的0%显著提升,虽操作量增加3.3倍,但净效率仍提高17%-18%。统计分析显示差异高度显著(p < 0.001,Cohen's d > 23),首次为大模型中主动记忆优势提供量化证据。本文整合50余篇最新研究,构建理论框架,推动从信息检索迈向真正的认知增强。
原文摘要 · Abstract (English)
Large Language Models (LLMs) face fundamental limitations in context management despite recent advances extending context windows to millions of tokens. We propose Cognitive Workspace, a novel paradigm that transcends traditional Retrieval-Augmented Generation (RAG) by emulating human cognitive mechanisms of external memory use. Drawing from cognitive science foundations including Baddeley's working memory model, Clark's extended mind thesis, and Hutchins' distributed cognition framework, we demonstrate that current passive retrieval systems fail to capture the dynamic, task-driven nature of human memory management. Our analysis of 2024-2025 developments reveals that while techniques like Infini-attention and StreamingLLM achieve impressive context lengths, they lack the metacognitive awareness and active planning capabilities essential for true cognitive extension. Cognitive Workspace addresses these limitations through three core innovations: (1) active memory management with deliberate information curation, (2) hierarchical cognitive buffers enabling persistent working states, and (3) task-driven context optimization that dynamically adapts to cognitive demands. Empirical validation demonstrates Cognitive Workspace achieves an average 58.6% memory reuse rate (ranging from 54-60% across different tasks) compared to 0% for traditional RAG, with 17-18% net efficiency gain despite 3.3x higher operation counts. Statistical analysis confirms these advantages with p < 0.001 and Cohen's d > 23 across multiple task types, establishing the first quantitative evidence for active memory superiority in LLM systems. We present a comprehensive theoretical framework synthesizing insights from 50+ recent papers, positioning Cognitive Workspace as a fundamental shift from information retrieval to genuine cognitive augmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。