用互动小说测试大模型长期记忆与多轮推理能力
StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns
- 基于动态分支剧情构建多轮交互环境
- 可区分即时反馈与自主回溯两种推理模式
- 适合评估模型在复杂场景下的记忆保持能力
长期记忆(LTM)对大语言模型在复杂演化环境中实现自主智能至关重要。尽管已有大量研究聚焦于增强记忆或基于检索的架构,但缺乏标准化基准来系统评估模型的长期记忆能力。现有基准在知识保留和动态序列推理评估方面仍存在不足,且灵活性有限,难以全面检验模型的LTM表现。为此,我们提出一个基于互动小说游戏的新基准框架,包含具有复杂推理结构的动态分支剧情,模拟真实世界场景,要求模型在层级决策树中导航,每项选择均引发多轮交互中的级联依赖。该基准强调两种不同推理设置:一种是错误决策后立即反馈,另一种则需模型自主追溯并修正先前选择。作为基准组成部分,我们还构建了一个用于测试模型在叙事驱动环境下长期记忆的新数据集。通过详尽实验验证,结果表明该基准能稳健、可靠地评估大模型的长期记忆能力。
原文摘要 · Abstract (English)
Long-term memory (LTM) is essential for large language models (LLMs) to achieve autonomous intelligence in complex, evolving environments. Despite increasing efforts in memory-augmented and retrieval-based architectures, there remains a lack of standardized benchmarks to systematically evaluate LLMs' long-term memory abilities. Existing benchmarks still face challenges in evaluating knowledge retention and dynamic sequential reasoning, and in their own flexibility, all of which limit their effectiveness in assessing models' LTM capabilities. To address these gaps, we propose a novel benchmark framework based on interactive fiction games, featuring dynamically branching storylines with complex reasoning structures. These structures simulate real-world scenarios by requiring LLMs to navigate hierarchical decision trees, where each choice triggers cascading dependencies across multi-turn interactions. Our benchmark emphasizes two distinct settings to test reasoning complexity: one with immediate feedback upon incorrect decisions, and the other requiring models to independently trace back and revise earlier choices after failure. As part of this benchmark, we also construct a new dataset designed to test LLMs' LTM within narrative-driven environments. We further validate the effectiveness of our approach through detailed experiments. Experimental results demonstrate the benchmark's ability to robustly and reliably assess LTM in LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。