arXiv:2603.12572cs.CL2026-03被引 7

构建首个长时记忆嵌入评估基准,测试模型在复杂记忆任务中的表现。

LMEB: Long-horizon Memory Embedding Benchmark

  • 设计22个数据集与193个零样本任务,覆盖四类记忆类型。
  • 大模型不总更优,传统检索能力与长时记忆无关。
  • 适合研究长期记忆、上下文依赖检索的学者使用。

记忆嵌入对记忆增强系统(如OpenClaw)至关重要,但现有文本嵌入基准仅聚焦于传统段落检索,未能评估模型处理碎片化、依赖上下文且时间跨度大的长时记忆检索任务的能力。为此,我们提出长时记忆嵌入基准(LMEB),一个全面的评估框架,用于衡量嵌入模型在复杂长时记忆检索中的表现。LMEB包含22个数据集和193个零样本检索任务,涵盖四种记忆类型:情景记忆、对话记忆、语义记忆和程序性记忆。这些类型在抽象层次和时间依赖性上各不相同,反映了现实世界中记忆检索的多样挑战。我们评估了15种广泛使用的嵌入模型,参数量从数亿到上百亿不等。结果表明:(1) LMEB具有合理难度;(2) 模型越大并不一定表现越好;(3) LMEB与MTEB衡量的是正交能力。这说明当前尚无通用模型能在所有记忆检索任务中表现优异,且传统段落检索能力强不代表长时记忆检索也强。LMEB提供了一个标准化、可复现的评估框架,填补了记忆嵌入评估的关键空白,支持未来长期、上下文依赖检索的发展。

原文摘要 · Abstract (English)

Memory embeddings are crucial for memory-augmented systems, such as OpenClaw, but their evaluation is underexplored in current text embedding benchmarks, which narrowly focus on traditional passage retrieval and fail to assess models' ability to handle long-horizon memory retrieval tasks involving fragmented, context-dependent, and temporally distant information. To address this gap, we introduce the Long-horizon Memory Embedding Benchmark (LMEB), a comprehensive framework for evaluating embedding models on complex, long-horizon memory retrieval. LMEB comprises 22 datasets and 193 zero-shot retrieval tasks spanning four memory types: episodic, dialogue, semantic, and procedural. These memory types differ in terms of level of abstraction and temporal dependency, capturing distinct aspects of memory retrieval that reflect the diverse challenges of the real world. We evaluate 15 widely used embedding models, ranging from hundreds of millions to ten billion parameters. The results reveal that (1) LMEB provides a reasonable level of difficulty; (2) Larger models do not always perform better; (3) LMEB and MTEB measure orthogonal capabilities. This suggests that the field has yet to converge on a universal model capable of excelling across all memory retrieval tasks, and that strong performance on traditional passage retrieval does not necessarily transfer to long-horizon memory retrieval. LMEB provides a standardized and reproducible framework that fills a key gap in memory embedding evaluation and supports future advances in long-term, context-dependent retrieval.

记忆嵌入长时记忆评估基准检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。