测试大模型能否像角色一样记住自己的故事,发现非参数方法更擅长持续学习。
If an LLM Were a Character, Would It Know Its Own Story? Evaluating Lifelong Learning in LLMs
- 构建剧本式数据集,模拟角色在多轮对话中积累记忆
- 非参数方法在追踪角色状态和关系上表现远超传统参数化模型
- 所有模型都出现记忆丢失问题,提示长期学习仍需突破
大型语言模型(LLMs)能进行类人对话,但因无状态特性而缺乏持续记忆。然而,在多轮、多智能体交互中,模型开始表现出一致的角色行为,暗示涌现的终身学习能力。现有评估基准多聚焦静态开放问答,难以捕捉此类动态。为此,我们提出 LIFESTATE-BENCH,一个基于《哈姆雷特》和合成剧本的剧集式数据集,富含叙事结构与角色互动。通过事实核查任务,评估模型对自我认知、情节记忆和关系追踪的能力,涵盖参数与非参数方法。在 Llama3.1-8B、GPT-4-turbo、DeepSeek R1 等模型上的实验表明,非参数方法显著优于参数方法;但所有模型在长时间交互中均出现灾难性遗忘,凸显终身学习仍需重大改进。
原文摘要 · Abstract (English)
Large language models (LLMs) can carry out human-like dialogue, but unlike humans, they are stateless due to the superposition property. However, during multi-turn, multi-agent interactions, LLMs begin to exhibit consistent, character-like behaviors, hinting at a form of emergent lifelong learning. Despite this, existing benchmarks often fail to capture these dynamics, primarily focusing on static, open-ended evaluations. To address this gap, we introduce LIFESTATE-BENCH, a benchmark designed to assess lifelong learning in LLMs. It features two episodic datasets: Hamlet and a synthetic script collection, rich in narrative structure and character interactions. Our fact checking evaluation probes models' self-awareness, episodic memory retrieval, and relationship tracking, across both parametric and non-parametric approaches. Experiments on models like Llama3.1-8B, GPT-4-turbo, and DeepSeek R1, we demonstrate that nonparametric methods significantly outperform parametric ones in managing stateful learning. However, all models exhibit challenges with catastrophic forgetting as interactions extend, highlighting the need for further advancements in lifelong learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。