为大模型设计动态记忆架构,提升长对话连贯性。
Memory-Augmented Architecture for Long-Term Context Handling in Large Language Models

- 动态检索、更新、清理历史交互信息
- 显著提升上下文连贯性,降低内存开销
- 适合需要长期对话的实时交互系统
大语言模型在长时间对话中因上下文记忆有限,常出现交流断裂和回复相关性下降,影响用户体验。为此,我们提出一种增强记忆的架构,可动态从过往交互中检索、更新并修剪相关信息,实现有效的长期上下文管理。实验表明,该方法显著提升上下文连贯性,减少内存占用,并改善回复质量,展现出在实时交互系统中的应用潜力。
原文摘要 · Abstract (English)
Large Language Models face significant challenges in maintaining coherent interactions over extended dialogues due to their limited contextual memory. This limitation often leads to fragmented exchanges and reduced relevance in responses, diminishing user experience. To address these issues, we propose a memory-augmented architecture that dynamically retrieves, updates, and prunes relevant information from past interactions, ensuring effective long-term context handling. Experimental results demonstrate that our solution significantly improves contextual coherence, reduces memory overhead, and enhances response quality, showcasing its potential for real-time applications in interactive systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。