为大模型设计记忆能力评测基准,揭示其在事件回忆上的短板。
Episodic Memories Generation and Evaluation Benchmark for Large Language Models
- 构建基于认知科学的事件记忆表示框架,包含时空与实体信息。
- 在10万词内文本中,顶尖模型仍难处理多事件关联与复杂时空关系。
- 开源数据集与代码,支持记忆生成与推理任务评估,适合认知智能研究者。
情景记忆——即对时间与空间上具体事件的回忆能力——是人类认知的核心,支撑连贯叙事、规划与决策。尽管大型语言模型(LLMs)能力卓越,却缺乏稳健的情景记忆机制。我们主张将情景记忆融入LLM,是推动AI向类人认知迈进的关键,有助于实现一致推理并使输出扎根于真实事件,从而减少虚构内容。为此,我们提出一个全面的框架,用于建模与评估LLM的情景记忆能力。受认知科学启发,我们设计了结构化方法来表征情景事件,涵盖时间空间背景、参与实体及详细描述。我们构建了一个无污染的独特情景记忆基准,公开发布源代码与数据集,用于评估模型在多种回忆与情景推理任务中的表现。对GPT-4、Claude系列、Llama 3.1和o1-mini等前沿模型的评估显示,即使最先进的模型在处理多个相关事件或复杂时空关系时仍表现不佳,即便在10k–100k token上下文中亦然。
原文摘要 · Abstract (English)
Episodic memory -- the ability to recall specific events grounded in time and space -- is a cornerstone of human cognition, enabling not only coherent storytelling, but also planning and decision-making. Despite their remarkable capabilities, Large Language Models (LLMs) lack a robust mechanism for episodic memory: we argue that integrating episodic memory capabilities into LLM is essential for advancing AI towards human-like cognition, increasing their potential to reason consistently and ground their output in real-world episodic events, hence avoiding confabulations. To address this challenge, we introduce a comprehensive framework to model and evaluate LLM episodic memory capabilities. Drawing inspiration from cognitive science, we develop a structured approach to represent episodic events, encapsulating temporal and spatial contexts, involved entities, and detailed descriptions. We synthesize a unique episodic memory benchmark, free from contamination, and release open source code and datasets to assess LLM performance across various recall and episodic reasoning tasks. Our evaluation of state-of-the-art models, including GPT-4 and Claude variants, Llama 3.1, and o1-mini, reveals that even the most advanced LLMs struggle with episodic memory tasks, particularly when dealing with multiple related events or complex spatio-temporal relationships -- even in contexts as short as 10k-100k tokens.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。