arXiv:2505.22921cs.CL2025-05被引 12

为大模型设计稳定记忆机制,提升长文本理解与对话连贯性。

Structured Memory Mechanisms for Stable Context Representation in Large Language Models

  • 引入显式记忆单元与门控写入,动态更新记忆内容。
  • 在多轮问答和跨段推理中显著减少语义漂移,提升生成一致性。
  • 适合需要长期上下文理解的复杂对话与长文本任务。

本文针对大语言模型在理解长期上下文时的局限性,提出一种具备长期记忆机制的模型架构,以增强段落间和对话回合间语义信息的保留与检索能力。该模型集成显式记忆单元、门控写入机制和基于注意力的读取模块,并引入遗忘函数实现记忆内容的动态更新,提升历史信息管理能力。为优化记忆操作效果,研究设计了联合训练目标,将主任务损失与记忆写入和遗忘的约束结合,引导模型学习更优的记忆策略。系统评估显示,该模型在文本生成一致性、多轮问答稳定性以及跨上下文推理准确性方面均表现显著优势。尤其在长文本任务和复杂问答场景中,展现出强语义保留与上下文连贯性,有效缓解传统模型处理长期依赖时的上下文丢失与语义漂移问题。实验还分析了不同记忆结构、容量大小和控制策略,进一步验证了记忆机制在语言理解中的关键作用,证实了所提方法在架构设计与性能表现上的可行性与有效性。

原文摘要 · Abstract (English)

This paper addresses the limitations of large language models in understanding long-term context. It proposes a model architecture equipped with a long-term memory mechanism to improve the retention and retrieval of semantic information across paragraphs and dialogue turns. The model integrates explicit memory units, gated writing mechanisms, and attention-based reading modules. A forgetting function is introduced to enable dynamic updates of memory content, enhancing the model's ability to manage historical information. To further improve the effectiveness of memory operations, the study designs a joint training objective. This combines the main task loss with constraints on memory writing and forgetting. It guides the model to learn better memory strategies during task execution. Systematic evaluation across multiple subtasks shows that the model achieves clear advantages in text generation consistency, stability in multi-turn question answering, and accuracy in cross-context reasoning. In particular, the model demonstrates strong semantic retention and contextual coherence in long-text tasks and complex question answering scenarios. It effectively mitigates the context loss and semantic drift problems commonly faced by traditional language models when handling long-term dependencies. The experiments also include analysis of different memory structures, capacity sizes, and control strategies. These results further confirm the critical role of memory mechanisms in language understanding. They demonstrate the feasibility and effectiveness of the proposed approach in both architectural design and performance outcomes.

长上下文记忆机制大模型对话系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。