为大模型长时记忆设计游戏化评估框架,更真实测试其持续记忆能力。
MemGround: Long-Term Memory Evaluation Kit for Large Language Models in Gamified Scenarios

- 构建三层次记忆评估体系,覆盖状态、时间关联与推理记忆
- 提出多维指标,量化记忆使用与行为轨迹,揭示模型短板
- 在游戏化交互中发现当前模型难以维持动态跟踪和复杂推理
现有大模型长时记忆评估仍以静态为主,聚焦简单检索与短上下文推理,忽视了连续交互中动态状态追踪与分层推理等复杂记忆特性。为此,我们提出MemGround——一个基于丰富游戏化交互场景的严谨长时记忆基准。通过引入三级分层框架,系统评估表面状态记忆、时间关联记忆与基于推理的记忆,并设计多维度度量体系:问答总分(QA Overall)、解锁记忆片段数(MFU)、正确顺序记忆片段数(MFCO)及探索轨迹图(ETD)。大量实验表明,当前顶尖大模型与记忆代理在交互环境中仍难以实现持续动态追踪、时间事件关联及基于长期积累证据的复杂推理。
原文摘要 · Abstract (English)
Current evaluations of long-term memory in LLMs are fundamentally static. By fixating on simple retrieval and short-context inference, they neglect the multifaceted nature of complex memory systems, such as dynamic state tracking and hierarchical reasoning in continuous interactions. To overcome these limitations, we propose MemGround, a rigorous long-term memory benchmark natively grounded in rich, gamified interactive scenarios. To systematically assess these capabilities, MemGround introduces a three-tier hierarchical framework that evaluates Surface State Memory, Temporal Associative Memory, and Reasoning-Based Memory through specialized interactive tasks. Furthermore, to comprehensively quantify both memory utilization and behavioral trajectories, we propose a multi-dimensional metric suite comprising Question-Answer Score (QA Overall), Memory Fragments Unlocked (MFU), Memory Fragments with Correct Order (MFCO), and Exploration Trajectory Diagrams (ETD). Extensive experiments reveal that state-of-the-art LLMs and memory agents still struggle with sustained dynamic tracking, temporal event association, and complex reasoning derived from long-term accumulated evidence in interactive environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。