arXiv:2511.03506cs.CL2025-11被引 43

首个针对记忆系统幻觉的分阶段评估基准,揭示记忆提取与更新中的错误生成。

HaluMem: Evaluating Hallucinations in Memory Systems of Agents

  • 设计三类任务分阶段检测记忆系统的幻觉行为
  • 在超长对话中发现记忆更新阶段易产生并累积错误
  • 适合关注大模型记忆可靠性与鲁棒性的研究者

记忆系统是实现大语言模型和智能体长期学习与持续交互的关键组件。然而,在存储与检索过程中,这些系统常出现幻觉,包括编造、错误、冲突与遗漏。现有评估多采用端到端问答,难以定位幻觉发生的具体操作阶段。为此,我们提出首个面向记忆系统操作层级的幻觉评估基准HaluMem,定义记忆提取、更新与问答三类任务,全面揭示不同阶段的幻觉表现。构建了以用户为中心的多轮人机交互数据集HaluMem-Medium和HaluMem-Long,每套包含约1.5万条记忆点与3500个异构问题,单用户平均对话长度达1500和2600轮,上下文长度超100万标记符,支持跨规模与复杂度的幻觉评估。实证研究表明,现有记忆系统在提取与更新阶段倾向于生成并累积幻觉,进而传播至问答阶段。未来研究应聚焦可解释、受控的记忆操作机制,系统性抑制幻觉,提升记忆可靠性。

原文摘要 · Abstract (English)

Memory systems are key components that enable AI systems such as LLMs and AI agents to achieve long-term learning and sustained interaction. However, during memory storage and retrieval, these systems frequently exhibit memory hallucinations, including fabrication, errors, conflicts, and omissions. Existing evaluations of memory hallucinations are primarily end-to-end question answering, which makes it difficult to localize the operational stage within the memory system where hallucinations arise. To address this, we introduce the Hallucination in Memory Benchmark (HaluMem), the first operation level hallucination evaluation benchmark tailored to memory systems. HaluMem defines three evaluation tasks (memory extraction, memory updating, and memory question answering) to comprehensively reveal hallucination behaviors across different operational stages of interaction. To support evaluation, we construct user-centric, multi-turn human-AI interaction datasets, HaluMem-Medium and HaluMem-Long. Both include about 15k memory points and 3.5k multi-type questions. The average dialogue length per user reaches 1.5k and 2.6k turns, with context lengths exceeding 1M tokens, enabling evaluation of hallucinations across different context scales and task complexities. Empirical studies based on HaluMem show that existing memory systems tend to generate and accumulate hallucinations during the extraction and updating stages, which subsequently propagate errors to the question answering stage. Future research should focus on developing interpretable and constrained memory operation mechanisms that systematically suppress hallucinations and improve memory reliability.

记忆系统幻觉评估大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。