构建MemBench评测集,全面评估大模型代理的记忆能力。
MemBench: Towards More Comprehensive Evaluation on the Memory of LLM-based Agents
- 区分事实记忆与反思记忆,设计参与和观察两种交互场景。
- 从有效性、效率、容量三方面评估,覆盖多维度记忆表现。
- 开源数据集与代码,助力记忆机制研究进展。
近期研究强调了大模型代理中记忆机制的重要性,使其能存储观测信息并适应动态环境。然而,当前对记忆能力的评估仍面临挑战:以往评测在记忆层次多样性与交互场景覆盖上不足,且缺乏多维度评估指标。为此,本文构建了一个更全面的数据集与基准评测体系。数据集包含事实记忆与反思记忆两类层次,并设计参与与观察两种交互场景。基于此,提出名为MemBench的评测基准,从有效性、效率与容量三个维度综合评估大模型代理的记忆能力。为促进研究发展,我们已将数据集与项目开源至https://github.com/import-myself/Membench。
原文摘要 · Abstract (English)
Recent works have highlighted the significance of memory mechanisms in LLM-based agents, which enable them to store observed information and adapt to dynamic environments. However, evaluating their memory capabilities still remains challenges. Previous evaluations are commonly limited by the diversity of memory levels and interactive scenarios. They also lack comprehensive metrics to reflect the memory capabilities from multiple aspects. To address these problems, in this paper, we construct a more comprehensive dataset and benchmark to evaluate the memory capability of LLM-based agents. Our dataset incorporates factual memory and reflective memory as different levels, and proposes participation and observation as various interactive scenarios. Based on our dataset, we present a benchmark, named MemBench, to evaluate the memory capability of LLM-based agents from multiple aspects, including their effectiveness, efficiency, and capacity. To benefit the research community, we release our dataset and project at https://github.com/import-myself/Membench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。