让大模型学会记忆时间线索的对话经历,更像真人陪伴。
Echo: A Large Language Model with Temporal Episodic Memory
- 用多智能体生成带时间线索的对话数据,训练模型记住事件细节。
- 在跨时长、多轮对话测试中,性能远超现有大模型。
- 适合做情感陪伴、个人助理等需长期记忆的AI应用。
大型语言模型在数学、编程和文学创作等领域表现卓越,但多数研究聚焦于基于语义记忆的问答,忽视了其处理情景记忆(EM)相关问题的潜力。这导致在情感陪伴、个人智能助手和AI教师等需情景记忆的应用中表现不佳。为此,我们提出Echo——一个具备时间情景记忆能力的大模型。通过多智能体数据生成框架,构建了复杂多轮情景记忆对话数据集(EM-Train),并将时间信息创新性融入训练过程。同时,我们设计了专门评估情景记忆能力的EM-Test基准,涵盖不同时间跨度与难度,全面评估多轮情景记忆对话表现。实验表明,Echo在EM-Test上显著优于现有先进模型。定性分析显示,其具备类人情景记忆能力。所有数据集、代码及模型权重将开源。
原文摘要 · Abstract (English)
Research on large language models (LLMs) has shown remarkable performance in domains such as mathematics, programming, and literary creation. However, most studies have focused on semantic memory-based question answering, neglecting LLMs' potential to handle episodic memory (EM)-related queries. This oversight has led to suboptimal performance in applications requiring EM, including emotional companionship, personal AI assistants, and AI teachers. To address this gap, we introduce Echo, a LLM enhanced with temporal episodic memory. We propose a Multi-Agent Data Generation Framework that guides the model in generating multi-turn, complex scenario episodic memory dialogue data (EM-Train). Temporal information is innovatively incorporated into the LLM training process, and Echo is trained using the EM-Train. Furthermore, We develop an EM-Test benchmark specifically designed to evaluate LLMs' episodic memory capabilities. The EM-Test assesses performance across various time spans and difficulty levels, providing a comprehensive evaluation of multi-turn episodic memory dialogues. Our experiments demonstrate that Echo significantly outperforms state-of-the-art LLMs on EM-Test. Additionally, a qualitative analysis reveals Echo's potential to exhibit human-like episodic memory capabilities. We will open-source all datasets, code, and model weights.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。